RAL Tier1 Operations Report for 11th January 2017
Review of Issues during the week 4th to 11th January 2017.
|
- We have been seeing SAM SRM tests failures for CMS. These are owing to the total load. On Friday an adjustment was made to the number of transfers Castor allocates to the newer three disk servers - which may help but not resolve the problem.
Resolved Disk Server Issues
|
Current operational status and issues
|
- There is a problem seen by LHCb of a low but persistent rate of failure when copying the results of batch jobs to Castor. There is also a further problem that sometimes occurs when these (failed) writes are attempted to storage at other sites.
- LHCb have reported a problem accessing some files - and a GGUS ticket is open about this.
Ongoing Disk Server Issues
|
- GDSS665 (LhcbRawRdst - D0T1) failed on Saturday 31st Dec. Two disks in the system were replaced and it was returned to service on Friday 6th Jan.
- GDSS780 (LHCbDst - D1T0) failed on Thursday 5th Jan. It was returned to service the following day - initially in read-only mode. The BIOS and IPMI firmware were updated.
Notable Changes made since the last meeting.
|
Service
|
Scheduled?
|
Outage/At Risk
|
Start
|
End
|
Duration
|
Reason
|
Castor CMS instance
|
SCHEDULED
|
OUTAGE
|
31/01/2017 10:00
|
31/01/2017 16:00
|
6 hours
|
Castor 2.1.15 Upgrade. Only affecting CMS instance. (CMS stager component being upgraded).
|
Castor GEN instance
|
SCHEDULED
|
OUTAGE
|
26/01/2017 10:00
|
26/01/2017 16:00
|
6 hours
|
Castor 2.1.15 Upgrade. Only affecting GEN instance. (GEN stager component being upgraded).
|
Castor Atlas instance
|
SCHEDULED
|
OUTAGE
|
24/01/2017 10:00
|
24/01/2017 16:00
|
6 hours
|
Castor 2.1.15 Upgrade. Only affecting Atlas instance. (Atlas stager component being upgraded).
|
Castor LHCb instance
|
SCHEDULED
|
OUTAGE
|
17/01/2017 10:00
|
17/01/2017 16:00
|
6 hours
|
Castor 2.1.15 Upgrade. Only affecting LHCb instance. (LHCb stager component being upgraded).
|
Advanced warning for other interventions
|
The following items are being discussed and are still to be formally scheduled and announced.
|
Pending - but not yet formally announced:
- Merge AtlasScratchDisk into larger Atlas disk pool.
Listing by category:
- Castor:
- Update to Castor version 2.1.15. Dates announced via GOC DB for early 2017.
- Update SRMs to new version, including updating to SL6. This will be done after the Castor 2.1.15 update.
- Fabric
- Firmware updates on older disk servers.
Entries in GOC DB starting since the last report.
|
Service
|
Scheduled?
|
Outage/At Risk
|
Start
|
End
|
Duration
|
Reason
|
All Castor storage (All SRMs)
|
SCHEDULED
|
OUTAGE
|
10/01/2017 10:00
|
10/01/2017 14:03
|
4 hours and 3 minutes
|
Castor 2.1.15 Upgrade. Upgrade of Nameserver component. All instances affected.
|
gridftp.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, xrootd.echo.stfc.ac.uk,
|
SCHEDULED
|
OUTAGE
|
05/01/2017 10:00
|
06/01/2017 15:00
|
1 day, 5 hours
|
ECHO Re-install
|
All Castor storage (All SRMs)
|
SCHEDULED
|
OUTAGE
|
05/01/2017 09:30
|
05/01/2017 11:54
|
2 hours and 24 minutes
|
Outage of Castor Storage System for patching
|
Open GGUS Tickets (Snapshot during morning of meeting)
|
GGUS ID |
Level |
Urgency |
State |
Creation |
Last Update |
VO |
Subject
|
125856
|
Green
|
Top Piority
|
In Progress
|
2017-01-06
|
2016-01-10
|
lhcB
|
Permission denied for some files
|
125480
|
Green
|
Less Urgent
|
On Hold
|
2016-12-09
|
2016-12-21
|
|
total Physical and Logical CPUs values
|
125157
|
Green
|
Less Urgent
|
In Progress
|
2016-11-24
|
2017-01-03
|
|
Creation of a repository within the EGI CVMFS infrastructure
|
124876
|
Amber
|
Less Urgent
|
On Hold
|
2016-11-07
|
2017-01-01
|
OPS
|
[Rod Dashboard] Issue detected : hr.srce.GridFTP-Transfer-ops@gridftp.echo.stfc.ac.uk
|
117683
|
Red
|
Less Urgent
|
On Hold
|
2015-11-18
|
2016-12-07
|
|
CASTOR at RAL not publishing GLUE 2. We looked at this as planned in December (report).
|
Key: Atlas HC = Atlas HammerCloud (Queue ANALY_RAL_SL6, Template 808); CMS HC = CMS HammerCloud
Day |
OPS |
Alice |
Atlas |
CMS |
LHCb |
Atlas HC |
CMS HC |
Comment
|
04/01/17 |
100 |
72 |
100 |
93 |
100 |
N/A |
100 |
Alice: Internal Alice problem; CMS: SRM test failures - User timeout
|
05/01/17 |
89.6 |
35 |
90 |
82 |
90 |
N/A |
100 |
Castor outage for patching. ; Alice: Internal Alice problem.; CMS: SRM test failures - User timeout
|
06/01/17 |
100 |
100 |
100 |
54 |
100 |
N/A |
100 |
SRM test failures - User timeout
|
07/01/17 |
100 |
100 |
100 |
95 |
100 |
N/A |
100 |
SRM test failures - User timeout
|
08/01/17 |
100 |
100 |
100 |
84 |
100 |
N/A |
100 |
SRM test failures - User timeout
|
09/01/17 |
100 |
100 |
100 |
95 |
100 |
N/A |
100 |
SRM test failures - User timeout
|
10/01/17 |
82.3 |
100 |
83 |
74 |
83 |
N/A |
N/A |
Castor upgrade (nameserver to 2.1.15). CMS: SRM test failures - User timeout
|
- It was reported that CMS are seeing poor CMS efficiencies across all Tier1s.
- We had to connect to Vidyo using the phone link. Will try and get Vidyo client fixed in our meeting room.