Difference between revisions of "Tier1 Operations Report 2017-01-11"

From GridPP Wiki
Jump to: navigation, search
(Created page with "==RAL Tier1 Operations Report for 11th January 2017== __NOTOC__ ====== ====== <!-- ************************************************************* -----> <!-- ***********Start...")
 
 
(14 intermediate revisions by one user not shown)
Line 10: Line 10:
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Review of Issues during the week 4th to 11th January 2017.
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Review of Issues during the week 4th to 11th January 2017.
 
|}
 
|}
Overall it was a fairly quiet holiday time operationally:
+
* We have been seeing SAM SRM tests failures for CMS. These are owing to the total load. On Friday an adjustment was made to the number of transfers Castor allocates to the newer three disk servers - which may help but not resolve the problem.
* There was a failure of a Power Distribution Unit in a rack in the UPS room in the early hours of Friday 23rd December. Staff attended on site to get the power back. This mainly affected internal services (monitoring etc) - but also the Top BDIIs. This was the second time this PDU had given problems and it was swapped out during Friday morning.  
+
* LHCb have reported a problem accessing some files - and a GGUS ticket is open about this.
* We have had load issues on the CMS Castor instance throughout the holiday period which has led to repeated SAM test failures.
+
 
<!-- ***********End Review of Issues during last week*********** ----->
 
<!-- ***********End Review of Issues during last week*********** ----->
 
<!-- *********************************************************** ----->
 
<!-- *********************************************************** ----->
Line 23: Line 22:
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Resolved Disk Server Issues
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Resolved Disk Server Issues
 
|}
 
|}
* GDSS671 (AliceDisk - D1T0) failed on Saturday (17th Dec) with a read-only filesystem. Three drives were replaced and the server was returned to service on the 22nd Dec.  
+
* GDSS665 (LhcbRawRdst - D0T1) failed on Saturday 31st Dec. Two disks in the system were replaced and it was returned to service on Friday 6th Jan.
 +
* GDSS780 (LHCbDst - D1T0) failed on Thursday 5th Jan. It was returned to service the following day - initially in read-only mode. The BIOS and IPMI firmware were updated.
 
<!-- ***************************************************** ----->
 
<!-- ***************************************************** ----->
  
Line 44: Line 44:
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Ongoing Disk Server Issues
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Ongoing Disk Server Issues
 
|}
 
|}
* GDSS665 (lhcbRawRdst - D0T1) failed on Saturday 31st Dec). Investigations are ongoing.
+
* None
 
<!-- ***************End Ongoing Disk Server Issues**************** ----->
 
<!-- ***************End Ongoing Disk Server Issues**************** ----->
 
<!-- ************************************************************* ----->
 
<!-- ************************************************************* ----->
Line 55: Line 55:
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notable Changes made since the last meeting.
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notable Changes made since the last meeting.
 
|}
 
|}
* None
+
* On Thursday 5th Jan the firmware was updated in the RAID cards in the Viglen '13 batch of disk servers.
 +
* cvmfs client v2.3.2-1 has been deployed on all worker nodes.
 +
* Increased the number of Castor transfer slots on the three newest disk servers in CMS_Tape.
 +
* Load balancers have been introduced in front of the Site-BDII systems.
 +
* The Castor Nameserver has been upgraded to version 2.1.15. This is the first step of the overall Castor 2.1.15 update.
 +
* Migration of LHCb data from 'C' to 'D' tapes ongoing. Now a little over 70% done. Around 280 out of the 1000 tapes still to do.
 
<!-- *************End Notable Changes made this last week************** ----->
 
<!-- *************End Notable Changes made this last week************** ----->
 
<!-- ****************************************************************** ----->
 
<!-- ****************************************************************** ----->
Line 107: Line 112:
 
| 6 hours
 
| 6 hours
 
| Castor 2.1.15 Upgrade. Only affecting LHCb instance. (LHCb stager component being upgraded).
 
| Castor 2.1.15 Upgrade. Only affecting LHCb instance. (LHCb stager component being upgraded).
|-
 
| All Castor (all SRM endpoints)
 
| SCHEDULED
 
| OUTAGE
 
| 10/01/2017 10:00
 
| 10/01/2017 16:00
 
| 6 hours
 
| Castor 2.1.15 Upgrade. Upgrade of Nameserver component. All instances affected.
 
|-
 
| gridftp.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, xrootd.echo.stfc.ac.uk,
 
| SCHEDULED
 
| OUTAGE
 
| 05/01/2017 10:00
 
| 06/01/2017 15:00
 
| 1 day, 5 hours
 
| ECHO Re-install
 
|-
 
| All Castor
 
| SCHEDULED
 
| OUTAGE
 
| 05/01/2017 09:30
 
| 05/01/2017 17:00
 
| 7 hours and 30 minutes
 
| Outage of Castor Storage System for patching
 
 
|}
 
|}
 
<!-- **********************End GOC DB Entries************************** ----->
 
<!-- **********************End GOC DB Entries************************** ----->
Line 173: Line 154:
 
! Reason
 
! Reason
 
|-
 
|-
|lcgbdii.gridpp.rl.ac.uk
+
| All Castor storage (All SRMs)
| UNSCHEDULED
+
| SCHEDULED
 
| OUTAGE
 
| OUTAGE
| 23/12/2016 01:30
+
| 10/01/2017 10:00
| 23/12/2016 03:30
+
| 10/01/2017 14:03
| 2 hours
+
| 4 hours and 3 minutes
| Networking problems affecting part of the Tier-1 service.
+
| Castor 2.1.15 Upgrade. Upgrade of Nameserver component. All instances affected.
 +
|-
 +
| gridftp.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, xrootd.echo.stfc.ac.uk,
 +
| SCHEDULED
 +
| OUTAGE
 +
| 05/01/2017 10:00
 +
| 06/01/2017 15:00
 +
| 1 day, 5 hours
 +
| ECHO Re-install
 +
|-
 +
| All Castor storage (All SRMs)
 +
| SCHEDULED
 +
| OUTAGE
 +
| 05/01/2017 09:30
 +
| 05/01/2017 11:54
 +
| 2 hours and 24 minutes
 +
| Outage of Castor Storage System for patching
 
|}
 
|}
 
<!-- **********************End GOC DB Entries************************** ----->
 
<!-- **********************End GOC DB Entries************************** ----->
Line 195: Line 192:
 
|-style="background:#b7f1ce"
 
|-style="background:#b7f1ce"
 
! GGUS ID !! Level !! Urgency !! State !! Creation !! Last Update !! VO !! Subject
 
! GGUS ID !! Level !! Urgency !! State !! Creation !! Last Update !! VO !! Subject
 +
|-
 +
| 125856
 +
| Green
 +
| Top Piority
 +
| In Progress
 +
| 2017-01-06
 +
| 2016-01-10
 +
| LHCb
 +
| Permission denied for some files
 
|-
 
|-
 
| 125480
 
| 125480
Line 215: Line 221:
 
|-
 
|-
 
| 124876
 
| 124876
| Yellow
+
| Amber
 
| Less Urgent
 
| Less Urgent
 
| On Hold
 
| On Hold
Line 248: Line 254:
 
! Day !! OPS !! Alice !! Atlas !! CMS !! LHCb !! Atlas HC !! CMS HC !! Comment
 
! Day !! OPS !! Alice !! Atlas !! CMS !! LHCb !! Atlas HC !! CMS HC !! Comment
 
|-
 
|-
| 21/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 96 || 100 || N/A || 100 || SRM test failures - User timeout
+
| 04/01/17 || 100 || style="background-color: lightgrey;" | 72 || 100 || style="background-color: lightgrey;" | 93 || 100 || N/A || 100 || Alice: Internal Alice problem; CMS: SRM test failures - User timeout
 
|-
 
|-
| 22/12/16 || 100 || 100 || 100 || 100 || 100 || N/A || 100 ||
+
| 05/01/17 || style="background-color: lightgrey;" | 89.6 || style="background-color: lightgrey;" | 35 || style="background-color: lightgrey;" | 90 || style="background-color: lightgrey;" | 82 || style="background-color: lightgrey;" | 90 || N/A || 100 || Castor outage for patching. ; Alice: Internal Alice problem.; CMS: SRM test failures - User timeout
|-
+
| 23/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 88 || 100 || N/A || N/A || SRM test failures - User timeout
+
|-
+
| 24/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 96 || 100 || N/A || 100 || SRM test failures - User timeout
+
|-
+
| 25/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 97 || 100 || N/A || 100 || SRM test failures - User timeout
+
|-
+
| 26/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 97 || 100 || N/A || 100 || SRM test failures - User timeout
+
|-
+
| 27/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 96 || 100 || N/A || N/A || SRM test failures - User timeout
+
|-
+
| 28/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 79 || 100 || N/A || 100 || SRM test failures - User timeout
+
|-
+
| 29/12/16 || 100 || 100 || style="background-color: lightgrey;" | 97 || style="background-color: lightgrey;" | 62 || 100 || N/A || 100 || CMS: Block of SRM test failures - User timeout; Atlas: single SRM test failure ( could not open connection to srm-atlas.gridpp.rl.ac.uk).
+
|-
+
| 30/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 95 || 100 || N/A || 100 || SRM test failures - User timeout
+
|-
+
| 31/12/16 || 100 || 100 || 100 || style="background-color: lightgrey;" | 97 || 100 || N/A || 100 || SRM test failures - User timeout
+
 
+
|-
+
| 04/01/17 || 100 || 72 || 100 || style="background-color: lightgrey;" | 93 || 100 || N/A || 100 || SRM test failures - User timeout
+
|-
+
| 05/01/17 || 100 || 100 || 100 || style="background-color: lightgrey;" | 82 || 100 || N/A || 100 || SRM test failures - User timeout
+
 
|-
 
|-
 
| 06/01/17 || 100 || 100 || 100 || style="background-color: lightgrey;" | 54  || 100 || N/A || 100 || SRM test failures - User timeout
 
| 06/01/17 || 100 || 100 || 100 || style="background-color: lightgrey;" | 54  || 100 || N/A || 100 || SRM test failures - User timeout
Line 283: Line 266:
 
| 09/01/17 || 100 || 100 || 100 || style="background-color: lightgrey;" | 95  || 100 || N/A || 100 || SRM test failures - User timeout
 
| 09/01/17 || 100 || 100 || 100 || style="background-color: lightgrey;" | 95  || 100 || N/A || 100 || SRM test failures - User timeout
 
|-
 
|-
| 10/01/17 || 100 || 100 || 100 || style="background-color: lightgrey;" | 74  || 100 || N/A || 100 || SRM test failures - User timeout
+
| 10/01/17 || style="background-color: lightgrey;" | 82.3 || 100 || style="background-color: lightgrey;" | 83 || style="background-color: lightgrey;" | 74  || style="background-color: lightgrey;" | 83 || N/A || N/A || Castor upgrade (nameserver to 2.1.15). CMS: SRM test failures - User timeout
 
+
 
|}
 
|}
 
<!-- **********************End Availability Report************************** ----->
 
<!-- **********************End Availability Report************************** ----->
Line 296: Line 278:
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notes from Meeting.
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notes from Meeting.
 
|}
 
|}
* It was reported that CMS are seeing poor CMS efficiencies across all Tier1s.
+
* The open GGUS tickets were reviewed.
* We had to connect to Vidyo using the phone link. Will try and get Vidyo client fixed in our meeting room.
+
* The ongoing CMS SRM AM test failures were discussed.
 +
* Following the successful Castor Nameserver update to version 2.1.15 the first of the stagers (LHCb) will be upgraded next week. The date for this has been announced as Tuesday (17th) - but owing to a change in staff availability it will be done on the Wednesday (18th).

Latest revision as of 18:01, 17 January 2017

RAL Tier1 Operations Report for 11th January 2017

Review of Issues during the week 4th to 11th January 2017.
  • We have been seeing SAM SRM tests failures for CMS. These are owing to the total load. On Friday an adjustment was made to the number of transfers Castor allocates to the newer three disk servers - which may help but not resolve the problem.
  • LHCb have reported a problem accessing some files - and a GGUS ticket is open about this.
Resolved Disk Server Issues
  • GDSS665 (LhcbRawRdst - D0T1) failed on Saturday 31st Dec. Two disks in the system were replaced and it was returned to service on Friday 6th Jan.
  • GDSS780 (LHCbDst - D1T0) failed on Thursday 5th Jan. It was returned to service the following day - initially in read-only mode. The BIOS and IPMI firmware were updated.
Current operational status and issues
  • There is a problem seen by LHCb of a low but persistent rate of failure when copying the results of batch jobs to Castor. There is also a further problem that sometimes occurs when these (failed) writes are attempted to storage at other sites.
Ongoing Disk Server Issues
  • None
Notable Changes made since the last meeting.
  • On Thursday 5th Jan the firmware was updated in the RAID cards in the Viglen '13 batch of disk servers.
  • cvmfs client v2.3.2-1 has been deployed on all worker nodes.
  • Increased the number of Castor transfer slots on the three newest disk servers in CMS_Tape.
  • Load balancers have been introduced in front of the Site-BDII systems.
  • The Castor Nameserver has been upgraded to version 2.1.15. This is the first step of the overall Castor 2.1.15 update.
  • Migration of LHCb data from 'C' to 'D' tapes ongoing. Now a little over 70% done. Around 280 out of the 1000 tapes still to do.
Declared in the GOC DB
Service Scheduled? Outage/At Risk Start End Duration Reason
Castor CMS instance SCHEDULED OUTAGE 31/01/2017 10:00 31/01/2017 16:00 6 hours Castor 2.1.15 Upgrade. Only affecting CMS instance. (CMS stager component being upgraded).
Castor GEN instance SCHEDULED OUTAGE 26/01/2017 10:00 26/01/2017 16:00 6 hours Castor 2.1.15 Upgrade. Only affecting GEN instance. (GEN stager component being upgraded).
Castor Atlas instance SCHEDULED OUTAGE 24/01/2017 10:00 24/01/2017 16:00 6 hours Castor 2.1.15 Upgrade. Only affecting Atlas instance. (Atlas stager component being upgraded).
Castor LHCb instance SCHEDULED OUTAGE 17/01/2017 10:00 17/01/2017 16:00 6 hours Castor 2.1.15 Upgrade. Only affecting LHCb instance. (LHCb stager component being upgraded).
Advanced warning for other interventions
The following items are being discussed and are still to be formally scheduled and announced.

Pending - but not yet formally announced:

  • Merge AtlasScratchDisk into larger Atlas disk pool.

Listing by category:

  • Castor:
    • Update to Castor version 2.1.15. Dates announced via GOC DB for early 2017.
    • Update SRMs to new version, including updating to SL6. This will be done after the Castor 2.1.15 update.
  • Fabric
    • Firmware updates on older disk servers.
Entries in GOC DB starting since the last report.
Service Scheduled? Outage/At Risk Start End Duration Reason
All Castor storage (All SRMs) SCHEDULED OUTAGE 10/01/2017 10:00 10/01/2017 14:03 4 hours and 3 minutes Castor 2.1.15 Upgrade. Upgrade of Nameserver component. All instances affected.
gridftp.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, xrootd.echo.stfc.ac.uk, SCHEDULED OUTAGE 05/01/2017 10:00 06/01/2017 15:00 1 day, 5 hours ECHO Re-install
All Castor storage (All SRMs) SCHEDULED OUTAGE 05/01/2017 09:30 05/01/2017 11:54 2 hours and 24 minutes Outage of Castor Storage System for patching
Open GGUS Tickets (Snapshot during morning of meeting)
GGUS ID Level Urgency State Creation Last Update VO Subject
125856 Green Top Piority In Progress 2017-01-06 2016-01-10 LHCb Permission denied for some files
125480 Green Less Urgent On Hold 2016-12-09 2016-12-21 total Physical and Logical CPUs values
125157 Green Less Urgent In Progress 2016-11-24 2017-01-03 Creation of a repository within the EGI CVMFS infrastructure
124876 Amber Less Urgent On Hold 2016-11-07 2017-01-01 OPS [Rod Dashboard] Issue detected : hr.srce.GridFTP-Transfer-ops@gridftp.echo.stfc.ac.uk
117683 Red Less Urgent On Hold 2015-11-18 2016-12-07 CASTOR at RAL not publishing GLUE 2. We looked at this as planned in December (report).
Availability Report

Key: Atlas HC = Atlas HammerCloud (Queue ANALY_RAL_SL6, Template 808); CMS HC = CMS HammerCloud

Day OPS Alice Atlas CMS LHCb Atlas HC CMS HC Comment
04/01/17 100 72 100 93 100 N/A 100 Alice: Internal Alice problem; CMS: SRM test failures - User timeout
05/01/17 89.6 35 90 82 90 N/A 100 Castor outage for patching. ; Alice: Internal Alice problem.; CMS: SRM test failures - User timeout
06/01/17 100 100 100 54 100 N/A 100 SRM test failures - User timeout
07/01/17 100 100 100 95 100 N/A 100 SRM test failures - User timeout
08/01/17 100 100 100 84 100 N/A 100 SRM test failures - User timeout
09/01/17 100 100 100 95 100 N/A 100 SRM test failures - User timeout
10/01/17 82.3 100 83 74 83 N/A N/A Castor upgrade (nameserver to 2.1.15). CMS: SRM test failures - User timeout
Notes from Meeting.
  • The open GGUS tickets were reviewed.
  • The ongoing CMS SRM AM test failures were discussed.
  • Following the successful Castor Nameserver update to version 2.1.15 the first of the stagers (LHCb) will be upgraded next week. The date for this has been announced as Tuesday (17th) - but owing to a change in staff availability it will be done on the Wednesday (18th).