Difference between revisions of "Tier1 Operations Report 2017-03-08"

From GridPP Wiki
Jump to: navigation, search
()
()
 
(22 intermediate revisions by one user not shown)
Line 10: Line 10:
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Review of Issues during the week 1st to 8th March 2017.
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Review of Issues during the week 1st to 8th March 2017.
 
|}
 
|}
* Castor:
+
* Following a discussion at last week's liaison meeting an apparant cap on LHCb batch jobs was removed. However, it turned out that the restriction in the Condor configuration file was not erroneous and had no effect. LHCb batch jobs were not being limited.
** There was a problem with the LHCb SRMs overnight last Wednesday - Thursday (23/24). This resolved itself in the morning.
+
* There was a problem reported for access to AtlasScratchDisk in Castor this morning (8th Mar). Atlas have reported a large backlog of outstanding file transfers. Being worked on at time of meeting.
** Atlas SRM SAM tests failed for several days. Our investigation suggest a problem with Atlas' test and we have raised a GGUS ticket with them.
+
** We still see some timeout test failures in SAM tests for CMS.
+
* Our investigations showed a higher than normal packet loss seen by the Perfsonar monitoring of our external data links. This started on the 14th February and disappeared on the 24th Feb. The cause is not known.
+
* CEPH ECHO: There was a problem with one 'placement group' that resulted in a loss of data (2000 Atlas files). This has been followed up and understood - in conjunction with the CEPH developers and is being presented in a CEPH forum. The understanding gained means that should this recur there would be no data loss.
+
* There was a problem with the CMS CE glexec tests this morning (1st March). This was found to be a problem with the argus server that was fixed around lunchtime. (This item added after the meeting).
+
 
<!-- ***********End Review of Issues during last week*********** ----->
 
<!-- ***********End Review of Issues during last week*********** ----->
 
<!-- *********************************************************** ----->
 
<!-- *********************************************************** ----->
Line 27: Line 22:
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Resolved Disk Server Issues
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Resolved Disk Server Issues
 
|}
 
|}
* GDSS663 (AtlasTape - D0T1) crashed on Saturday (18th Feb). Two faulty disks found and replaced. It was returned to service during the afternoon of Wednesday 22nd Feb.
+
* None
* GDSS662 (AtlasTape - D0T1) crashed in the early hours of Monday 27th Feb. It was returned to service the following day having had one disk drive replaced.
+
 
<!-- ***************************************************** ----->
 
<!-- ***************************************************** ----->
  
Line 49: Line 43:
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Ongoing Disk Server Issues
 
| style="background-color: #f8d6a9; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Ongoing Disk Server Issues
 
|}
 
|}
* None
+
* GDSS689 (AtlasDataDisk - D1T0) reported 'fsprobe' errors and was taken out of production this morning. Investigations ongoing.
 
<!-- ***************End Ongoing Disk Server Issues**************** ----->
 
<!-- ***************End Ongoing Disk Server Issues**************** ----->
 
<!-- ************************************************************* ----->
 
<!-- ************************************************************* ----->
Line 60: Line 54:
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Limits on concurrent batch system jobs.
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Limits on concurrent batch system jobs.
 
|}
 
|}
* LHCb Pilot 4500
 
 
* Atlas Pilot (Analysis) 1500
 
* Atlas Pilot (Analysis) 1500
 
* CMS Multicore 460
 
* CMS Multicore 460
Line 73: Line 66:
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notable Changes made since the last meeting.
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notable Changes made since the last meeting.
 
|}
 
|}
* Various systems have had security and other patches applied. More back-end database systems have been updated to remove a software layer ("asmlib").
+
* CMS PhEDEx debug transfers switched from CASTOR to CEPH ECHO.
* A successful UPS/Generator load test was carried out on Tuesday morning (28th Feb).
+
* Ongoing work applying security and other patches. More back-end database systems have been updated to remove a software layer ("asmlib").
 +
* IPv6 has been disabled across systems in preparation for enabling IPv6 through the routers.
 +
* This morning (8th March) work has been carried out to enabled IPv6 through the Tier's routers.
 
<!-- *************End Notable Changes made this last week************** ----->
 
<!-- *************End Notable Changes made this last week************** ----->
 
<!-- ****************************************************************** ----->
 
<!-- ****************************************************************** ----->
Line 85: Line 80:
 
| style="background-color: #d8e8ff; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Declared in the GOC DB
 
| style="background-color: #d8e8ff; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Declared in the GOC DB
 
|}
 
|}
{| border=1 align=center
+
None
|- bgcolor="#7c8aaf"
+
! Service
+
! Scheduled?
+
! Outage/At Risk
+
! Start
+
! End
+
! Duration
+
! Reason
+
|-
+
| Whole site
+
| SCHEDULED
+
| WARNING
+
| 08/03/2017 07:00
+
| 08/03/2017 11:00
+
| 4 hours
+
| Warning on site during network intervention in preparation for IPv6.
+
|}
+
 
<!-- **********************End GOC DB Entries************************** ----->
 
<!-- **********************End GOC DB Entries************************** ----->
 
<!-- ****************************************************************** ----->
 
<!-- ****************************************************************** ----->
Line 117: Line 95:
 
<!-- ******* still to be formally scheduled and/or announced ******* ----->
 
<!-- ******* still to be formally scheduled and/or announced ******* ----->
 
'''Pending - but not yet formally announced:'''
 
'''Pending - but not yet formally announced:'''
 +
* Update Castor SRMs. Propose LHCb SRMs first - target date 22nd March.
 +
* Chiller replacement - work imminent.
 
* Merge AtlasScratchDisk into larger Atlas disk pool.
 
* Merge AtlasScratchDisk into larger Atlas disk pool.
 
'''Listing by category:'''
 
'''Listing by category:'''
 
* Castor:  
 
* Castor:  
** Update SRMs to new version, including updating to SL6. This will be done after the Castor 2.1.15 update.
+
** Update SRMs to new version, including updating to SL6.
* Networking:
+
** Bring some newer disk servers ('14 generation) into service, replacing some older ('12 generation) servers.
** Enabling IPv6 onto production network.
+
 
* Databases
 
* Databases
** Removal of "asmlib" layer on Oracle database nodes.
+
** Removal of "asmlib" layer on Oracle database nodes. (Ongoing)
 +
* Infrastructure:
 +
** Two of the chillers supplying the air-conditioning for the R89 machine room will be replaced.
 
<!-- ***************End Advanced warning for other interventions*************** ----->
 
<!-- ***************End Advanced warning for other interventions*************** ----->
 
<!-- ************************************************************************** ----->
 
<!-- ************************************************************************** ----->
Line 145: Line 126:
 
! Reason
 
! Reason
 
|-
 
|-
| gridftp.echo.stfc.ac.uk, xrootd.echo.stfc.ac.uk,
+
| Whole site
| UNSCHEDULED
+
| OUTAGE
+
| 24/02/2017 15:00
+
| 24/02/2017 15:47
+
| 47 minutes
+
| Problem with ATLAS ECHO pool
+
|-
+
| lfc.gridpp.rl.ac.uk
+
 
| SCHEDULED
 
| SCHEDULED
 
| WARNING
 
| WARNING
| 22/02/2017 08:45
+
| 08/03/2017 07:00
| 22/02/2017 13:00
+
| 08/03/2017 11:00
| 4 hours and 15 minutes
+
| LFC Oracle backend security updates
+
|-
+
| All Castor and ECHO storage and Perfsonar.
+
| SCHEDULED
+
| WARNING
+
| 22/02/2017 07:00
+
| 22/02/2017 11:00
+
 
| 4 hours
 
| 4 hours
| Warning on Storage and Perfsonar during network intrevention in preparation for IPv6.
+
| Warning on site during network intervention in preparation for IPv6.
 
|}
 
|}
 
<!-- **********************End GOC DB Entries************************** ----->
 
<!-- **********************End GOC DB Entries************************** ----->
Line 184: Line 149:
 
! GGUS ID !! Level !! Urgency !! State !! Creation !! Last Update !! VO !! Subject
 
! GGUS ID !! Level !! Urgency !! State !! Creation !! Last Update !! VO !! Subject
 
|-
 
|-
| 126718
+
| 126905
 
| Green
 
| Green
| Urgent
+
| Less Urgent
 
| In Progress
 
| In Progress
| 2017-02-21
+
| 2017-03-02
| 2017-03-01
+
| 2017-03-02
| Atlas
+
| solid
| UK RAL-LCG2-ECHO DATADISK: ~8k deletion error due to "Device or resource busy"
+
| finish commissioning cvmfs server for solidexperiment.org
|-
+
| 126532
+
| Green
+
| Urgent
+
| In Progress
+
| 2017-02-09
+
| 2017-02-21
+
| Atlas
+
| RAL tape staging errors
+
 
|-
 
|-
 
| 126184
 
| 126184
Line 225: Line 181:
 
| On Hold
 
| On Hold
 
| 2015-11-18
 
| 2015-11-18
| 2017-02-10
+
| 2017-03-02
 
|  
 
|  
| CASTOR at RAL not publishing GLUE 2. Looking at it again now (Feb), progress made on back end. Need to update ticket.
+
| CASTOR at RAL not publishing GLUE 2.
 
|}
 
|}
 
<!-- **********************End GGUS Tickets************************** ----->
 
<!-- **********************End GGUS Tickets************************** ----->
Line 244: Line 200:
 
|-style="background:#b7f1ce"
 
|-style="background:#b7f1ce"
 
! Day !! OPS !! Alice !! Atlas !! CMS !! LHCb !! Atlas HC !! Atlas HC ECHO !! CMS HC !! Comment
 
! Day !! OPS !! Alice !! Atlas !! CMS !! LHCb !! Atlas HC !! Atlas HC ECHO !! CMS HC !! Comment
 +
|-
 
| 01/03/17 || 100 || 100 || style="background-color: lightgrey;" | 83 || style="background-color: lightgrey;" | 81 || 100 || 99 || 100 || 99 || Atlas: Ongoing problems with SRM test; CMS - CE test failures due to poblem with argus server.
 
| 01/03/17 || 100 || 100 || style="background-color: lightgrey;" | 83 || style="background-color: lightgrey;" | 81 || 100 || 99 || 100 || 99 || Atlas: Ongoing problems with SRM test; CMS - CE test failures due to poblem with argus server.
 
|-
 
|-
Line 256: Line 213:
 
| 06/03/17 || 100 || style="background-color: lightgrey;" | 68 || style="background-color: lightgrey;" | 89 || style="background-color: lightgrey;" | 97 || 100 || 99 || 98 || 100 || Alice: Central monitoring problem; Atlas: Ongoing problems with SRM test; CMS - timeouts in SRM tests.
 
| 06/03/17 || 100 || style="background-color: lightgrey;" | 68 || style="background-color: lightgrey;" | 89 || style="background-color: lightgrey;" | 97 || 100 || 99 || 98 || 100 || Alice: Central monitoring problem; Atlas: Ongoing problems with SRM test; CMS - timeouts in SRM tests.
 
|-
 
|-
| 07/03/17 || 100 || 100 || 100 || 100 || 100 || 100 || 100 || 100 || Timeouts on CMS SRM tests.
+
| 07/03/17 || 100 || 100 || 100 || style="background-color: lightgrey;" | 93 || 100 || 94 || 99 || 100 || Timeouts on CMS SRM tests.
 
|-
 
|-
 
|}
 
|}
Line 269: Line 226:
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notes from Meeting.
 
| style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Notes from Meeting.
 
|}
 
|}
* None yet
+
* ECHO: Two additional 'MON' boxes are being set-up bringing the total to five. The existing three can cope with normal activity but the additional ones would speed up recoveries and starts. Two additional gateway nodes are also being set-up (also bringing the total to five) which will improve bandwidth.
 +
* VO Dirac: some tar files moved from Edinburgh but little activity from sites other than Durham as yet.

Latest revision as of 15:14, 8 March 2017

RAL Tier1 Operations Report for 8th March 2017

Review of Issues during the week 1st to 8th March 2017.
  • Following a discussion at last week's liaison meeting an apparant cap on LHCb batch jobs was removed. However, it turned out that the restriction in the Condor configuration file was not erroneous and had no effect. LHCb batch jobs were not being limited.
  • There was a problem reported for access to AtlasScratchDisk in Castor this morning (8th Mar). Atlas have reported a large backlog of outstanding file transfers. Being worked on at time of meeting.
Resolved Disk Server Issues
  • None
Current operational status and issues
  • We are still seeing a rate of failures of the CMS SAM tests against the SRM. These are affecting our (CMS) availabilities but the level of failures is reduced as compared to a few weeks ago.
Ongoing Disk Server Issues
  • GDSS689 (AtlasDataDisk - D1T0) reported 'fsprobe' errors and was taken out of production this morning. Investigations ongoing.
Limits on concurrent batch system jobs.
  • Atlas Pilot (Analysis) 1500
  • CMS Multicore 460
Notable Changes made since the last meeting.
  • CMS PhEDEx debug transfers switched from CASTOR to CEPH ECHO.
  • Ongoing work applying security and other patches. More back-end database systems have been updated to remove a software layer ("asmlib").
  • IPv6 has been disabled across systems in preparation for enabling IPv6 through the routers.
  • This morning (8th March) work has been carried out to enabled IPv6 through the Tier's routers.
Declared in the GOC DB

None

Advanced warning for other interventions
The following items are being discussed and are still to be formally scheduled and announced.

Pending - but not yet formally announced:

  • Update Castor SRMs. Propose LHCb SRMs first - target date 22nd March.
  • Chiller replacement - work imminent.
  • Merge AtlasScratchDisk into larger Atlas disk pool.

Listing by category:

  • Castor:
    • Update SRMs to new version, including updating to SL6.
    • Bring some newer disk servers ('14 generation) into service, replacing some older ('12 generation) servers.
  • Databases
    • Removal of "asmlib" layer on Oracle database nodes. (Ongoing)
  • Infrastructure:
    • Two of the chillers supplying the air-conditioning for the R89 machine room will be replaced.
Entries in GOC DB starting since the last report.
Service Scheduled? Outage/At Risk Start End Duration Reason
Whole site SCHEDULED WARNING 08/03/2017 07:00 08/03/2017 11:00 4 hours Warning on site during network intervention in preparation for IPv6.
Open GGUS Tickets (Snapshot during morning of meeting)
GGUS ID Level Urgency State Creation Last Update VO Subject
126905 Green Less Urgent In Progress 2017-03-02 2017-03-02 solid finish commissioning cvmfs server for solidexperiment.org
126184 Green Less Urgent In Progress 2017-01-26 2017-02-07 Atlas Request of inputs for new sites monitoring
124876 Red Less Urgent On Hold 2016-11-07 2017-01-01 OPS [Rod Dashboard] Issue detected : hr.srce.GridFTP-Transfer-ops@gridftp.echo.stfc.ac.uk
117683 Red Less Urgent On Hold 2015-11-18 2017-03-02 CASTOR at RAL not publishing GLUE 2.
Availability Report

Key: Atlas HC = Atlas HammerCloud (Queue ANALY_RAL_SL6, Template 845); Atlas HC ECHO = Atlas ECHO (Template 842);CMS HC = CMS HammerCloud

Day OPS Alice Atlas CMS LHCb Atlas HC Atlas HC ECHO CMS HC Comment
01/03/17 100 100 83 81 100 99 100 99 Atlas: Ongoing problems with SRM test; CMS - CE test failures due to poblem with argus server.
02/03/17 100 100 59 99 100 100 99 100 Atlas: Ongoing problems with SRM test; CMS - timeouts in SRM tests.
03/03/17 100 100 91 98 100 97 100 100 Atlas: Ongoing problems with SRM test; CMS - timeouts in SRM tests.
04/03/17 100 100 96 100 100 100 100 99 Atlas: Ongoing problems with SRM test.
05/03/17 100 97 98 88 92 97 100 100 Atlas: Ongoing problems with SRM test; CMS - timeouts in SRM tests; LHCb - some SRM test failures.
06/03/17 100 68 89 97 100 99 98 100 Alice: Central monitoring problem; Atlas: Ongoing problems with SRM test; CMS - timeouts in SRM tests.
07/03/17 100 100 100 93 100 94 99 100 Timeouts on CMS SRM tests.
Notes from Meeting.
  • ECHO: Two additional 'MON' boxes are being set-up bringing the total to five. The existing three can cope with normal activity but the additional ones would speed up recoveries and starts. Two additional gateway nodes are also being set-up (also bringing the total to five) which will improve bandwidth.
  • VO Dirac: some tar files moved from Edinburgh but little activity from sites other than Durham as yet.