Tier1 Operations Report 2017-03-01

From GridPP Wiki
Jump to: navigation, search

RAL Tier1 Operations Report for 1st March 2017

Review of Issues during the week 22nd February to 1st March 2017.
  • Castor:
    • There was a problem with the LHCb SRMs overnight last Wednesday - Thursday (23/24). This resolved itself in the morning.
    • Atlas SRM SAM tests failed for several days. Our investigation suggest a problem with Atlas' test and we have raised a GGUS ticket with them.
    • We still see some timeout test failures in SAM tests for CMS.
  • Our investigations showed a higher than normal packet loss seen by the Perfsonar monitoring of our external data links. This started on the 14th February and disappeared on the 24th Feb. The cause is not known.
  • CEPH ECHO: There was a problem with one 'placement group' that resulted in a loss of data (2000 Atlas files). This has been followed up and understood - in conjunction with the CEPH developers and is being presented in a CEPH forum. The understanding gained means that should this recur there would be no data loss.
  • There was a problem with the CMS CE glexec tests this morning (1st March). This was found to be a problem with the argus server that was fixed around lunchtime. (This item added after the meeting).
Resolved Disk Server Issues
  • GDSS663 (AtlasTape - D0T1) crashed on Saturday (18th Feb). Two faulty disks found and replaced. It was returned to service during the afternoon of Wednesday 22nd Feb.
  • GDSS662 (AtlasTape - D0T1) crashed in the early hours of Monday 27th Feb. It was returned to service the following day having had one disk drive replaced.
Current operational status and issues
  • We are still seeing a rate of failures of the CMS SAM tests against the SRM. These are affecting our (CMS) availabilities but the level of failures is reduced as compared to a few weeks ago.
Ongoing Disk Server Issues
  • None
Limits on concurrent batch system jobs.
  • LHCb Pilot 4500
  • Atlas Pilot (Analysis) 1500
  • CMS Multicore 460
Notable Changes made since the last meeting.
  • Various systems have had security and other patches applied. More back-end database systems have been updated to remove a software layer ("asmlib").
  • A successful UPS/Generator load test was carried out on Tuesday morning (28th Feb).
Declared in the GOC DB
Service Scheduled? Outage/At Risk Start End Duration Reason
Whole site SCHEDULED WARNING 08/03/2017 07:00 08/03/2017 11:00 4 hours Warning on site during network intervention in preparation for IPv6.
Advanced warning for other interventions
The following items are being discussed and are still to be formally scheduled and announced.

Pending - but not yet formally announced:

  • Merge AtlasScratchDisk into larger Atlas disk pool.

Listing by category:

  • Castor:
    • Update SRMs to new version, including updating to SL6. This will be done after the Castor 2.1.15 update.
  • Networking:
    • Enabling IPv6 onto production network.
  • Databases
    • Removal of "asmlib" layer on Oracle database nodes.
Entries in GOC DB starting since the last report.
Service Scheduled? Outage/At Risk Start End Duration Reason
gridftp.echo.stfc.ac.uk, xrootd.echo.stfc.ac.uk, UNSCHEDULED OUTAGE 24/02/2017 15:00 24/02/2017 15:47 47 minutes Problem with ATLAS ECHO pool
lfc.gridpp.rl.ac.uk SCHEDULED WARNING 22/02/2017 08:45 22/02/2017 13:00 4 hours and 15 minutes LFC Oracle backend security updates
All Castor and ECHO storage and Perfsonar. SCHEDULED WARNING 22/02/2017 07:00 22/02/2017 11:00 4 hours Warning on Storage and Perfsonar during network intrevention in preparation for IPv6.
Open GGUS Tickets (Snapshot during morning of meeting)
GGUS ID Level Urgency State Creation Last Update VO Subject
126718 Green Urgent In Progress 2017-02-21 2017-03-01 Atlas UK RAL-LCG2-ECHO DATADISK: ~8k deletion error due to "Device or resource busy"
126532 Green Urgent In Progress 2017-02-09 2017-02-21 Atlas RAL tape staging errors
126184 Green Less Urgent In Progress 2017-01-26 2017-02-07 Atlas Request of inputs for new sites monitoring
124876 Red Less Urgent On Hold 2016-11-07 2017-01-01 OPS [Rod Dashboard] Issue detected : hr.srce.GridFTP-Transfer-ops@gridftp.echo.stfc.ac.uk
117683 Red Less Urgent On Hold 2015-11-18 2017-02-10 CASTOR at RAL not publishing GLUE 2. Looking at it again now (Feb), progress made on back end. Need to update ticket.
Availability Report

Key: Atlas HC = Atlas HammerCloud (Queue ANALY_RAL_SL6, Template 845); Atlas HC ECHO = Atlas ECHO (Template 842);CMS HC = CMS HammerCloud

Day OPS Alice Atlas CMS LHCb Atlas HC Atlas HC ECHO CMS HC Comment
22/02/17 100 100 98 95 100 97 96 100 Timeouts on CMS SRM tests.
23/02/17 100 100 100 95 100 99 100 100 Timeouts on CMS SRM tests.
24/02/17 100 100 75 97 100 100 99 100 Timeouts on CMS SRM tests.
25/02/17 100 100 18 98 100 98 100 100 Timeouts on CMS SRM tests.
26/02/17 100 100 34 98 100 100 100 100 Timeouts on CMS SRM tests.
27/02/17 100 100 2 95 100 97 100 99 Timeouts on CMS SRM tests.
28/02/17 100 100 97 95 100 98 100 99 Atlas: Previous problem stopped just after midnight; CMS: Timeouts on SRM tests.
Notes from Meeting.
  • Raja questioned the limit on LHCb pilot jobs (4500). This was removed shortly after the meeting.
  • There was a discussion around IPv6 plans. The enabling of IPv6 access to the Tier1 systems is expected to be completed in the next week or so. The first services that are planned to be made available over IPv6 are FTS and CVMFS Stratum 1.
  • Andrew reported a problem with cvmfs on SL7 that he has reported to the developers.