Difference between revisions of "Tier1 Operations Report 2018-05-28"

From GridPP Wiki
Jump to: navigation, search
()
Line 241: Line 241:
 
! Scope
 
! Scope
 
|-
 
|-
| 135164
+
| 135001
| none
+
| cms
| verified
+
| closed
| top priority
+
| urgent
| 16/05/2018
+
| 09/05/2018
| 16/05/2018
+
| 24/05/2018
| Other
+
| CMS_Data Transfers
| This TEST ALARM has been raised for testing GGUS alarm work flow after a new GGUS release.
+
| Fts-client needs to be updated
 
| WLCG
 
| WLCG
 
|-
 
|-
| 134853
+
| 134769
 
| cms
 
| cms
 
| closed
 
| closed
 
| urgent
 
| urgent
| 02/05/2018
+
| 26/04/2018
| 16/05/2018
+
| 22/05/2018
| CMS_Facilities
+
| CMS_Data Transfers
| T1_UK_RAL HammerCloud failures
+
| Transfers from RAL_Disk to Florida are failing
 
| WLCG
 
| WLCG
 
|-
 
|-
| 134839
+
| 134744
| snoplus.snolab.ca
+
| cms
 
| closed
 
| closed
| urgent
+
| top priority
| 30/04/2018
+
| 25/04/2018
| 16/05/2018
+
| 22/05/2018
 
| File Transfer
 
| File Transfer
| Data Transfer Failure
+
| Zero Phedex Transfers - via RAL FTS service on certain links
 
| EGI
 
| EGI
 
|-
 
|-
| 134737
+
| 134619
 
| cms
 
| cms
 
| closed
 
| closed
 
| urgent
 
| urgent
| 25/04/2018
+
| 19/04/2018
| 14/05/2018
+
| 22/05/2018
 
| CMS_SAM tests
 
| CMS_SAM tests
| SAM CE test failing at T1_UK_RAL
+
| Problems reading data from ECHO
| WLCG
+
|-
+
| 134494
+
| atlas
+
| closed
+
| urgent
+
| 11/04/2018
+
| 17/05/2018
+
| Storage Systems
+
| json space reporting not updated
+
 
| WLCG
 
| WLCG
 
|}
 
|}

Revision as of 12:31, 30 May 2018

RAL Tier1 Operations Report for 28th May 2018

Review of Issues during the week 14th May to the 21st May 2018.
  • No incidents(major or minor), have been flagged during this reporting period.
Current operational status and issues
  • None
Resolved Castor Disk Server Issues
  • gdss732 (lhcbDst- D1T0) - Back in production after completion of rebuilding of the replacement drive.
Ongoing Castor Disk Server Issues
  • None
Limits on concurrent batch system jobs.
  • CMS Multicore 550
Notable Changes made since the last meeting.
  • None.
Entries in GOC DB starting since the last report.
  • None
Declared in the GOC DB
  • None
Advanced warning for other interventions
The following items are being discussed and are still to be formally scheduled and announced.

Listing by category:

  • Castor:
    • Update systems to use SL7 and configured by Quattor/Aquilon. (Tape servers done)
    • Move to generic Castor headnodes.
  • Networking
    • Extend the number of services on the production network with IPv6 dual stack. (Done for Perfsonar, FTS3, all squids and the CVMFS Stratum-1 servers).
  • Internal
    • DNS servers will be rolled out within the Tier1 network.
  • Infrastructure
    • Testing of power distribution boards in the R89 machine room is being scheduled for some time late July / Early August. The effect of this on our services is being discussed.
Open

GGUS Tickets (Snapshot during morning of the report). The latest ticket snapshot can be found here[1].

Request id Affected vo Status Priority Date of creation Last update Type of problem Subject Scope
135133 cms in progress urgent 15/05/2018 17/05/2018 CMS_Data Transfers Likely corrupted File at T1_UK_RAL WLCG
134703 cms in progress urgent 23/04/2018 18/05/2018 CMS_Data Transfers Transfer failing from RAL_Disk WLCG
134685 dteam in progress less urgent 23/04/2018 02/05/2018 Middleware please upgrade perfsonar host(s) at RAL-LCG2 to CentOS7 EGI
134468 cms waiting for reply top priority 09/04/2018 18/05/2018 CMS_AAA WAN Access Xrootd redirector not seeing some files in ECHO WLCG
133992 atlas in progress less urgent 12/03/2018 19/04/2018 File Transfer RAL-LCG2-ECHO: No such file or directory EGI
127597 cms on hold urgent 07/04/2017 30/04/2018 File Transfer Check networking and xrootd RAL-CERN performance EGI
124876 ops on hold less urgent 07/11/2016 13/11/2017 Operations [Rod Dashboard] Issue detected : hr.srce.GridFTP-Transfer-ops@gridftp.echo.stfc.ac.uk EGI
117683 none on hold less urgent 18/11/2015 09/05/2018 Information System CASTOR at RAL not publishing GLUE 2 EGI
GGUS Tickets Closed Last week
Request id Affected vo Status Priority Date of creation Last update Type of problem Subject Scope
135001 cms closed urgent 09/05/2018 24/05/2018 CMS_Data Transfers Fts-client needs to be updated WLCG
134769 cms closed urgent 26/04/2018 22/05/2018 CMS_Data Transfers Transfers from RAL_Disk to Florida are failing WLCG
134744 cms closed top priority 25/04/2018 22/05/2018 File Transfer Zero Phedex Transfers - via RAL FTS service on certain links EGI
134619 cms closed urgent 19/04/2018 22/05/2018 CMS_SAM tests Problems reading data from ECHO WLCG
Availability Report
Target Availability for each site is 97.0% Red <90% Orange <97%
Day Atlas Atlas-Echo CMS LHCB Alice OPS Comments
2018-05-14 100 100 100 100 100 100
2018-05-15 100 100 99 100 96 100
2018-05-16 100 100 100 100 100 100
2018-05-17 100 100 98 100 100 100
2018-05-18 100 100 98 100 100 100
2018-05-19 100 100 96 100 100 100
2018-05-20 98 100 100 100 100 100
Hammercloud Test Report
Target Availability for each site is 97.0% Red <90% Orange <97%

Key: Atlas HC = Atlas HammerCloud (Queue RAL-LCG2_UCORE, Template 841); CMS HC = CMS HammerCloud

Day Atlas HC CMS HC Comment
2018/05/22 98 100
2018/05/23 98 98
2018/05/24 97 99
2018/05/25 96 99
2018/05/26 98 56
2018/05/27 100 60
2018/05/28 93 100
Notes from Meeting.
  • None yet