RAL Tier1 Operations Report for 22nd February 2017
Review of Issues during the week 15th to 22nd February 2017.
|
- Since the castor upgrade we have seen a couple of further problems which are now mainly fixed:
- Two VOs, Atlas and LHCb have seen issues with database resources (number of cursors). The latest being Atlas on Friday 10th. We still don't understand this.
- We had been failing tests for CMS xroot. This was fixed by changing the weighting for xrootd in the Castor transfermanagerd.
- We had being failing CMS tests for an SRM endpoint defined in the GOC DB but not in production ("srm-cms-disk"). This SRM endpoint was removed from the GOCDB but then we started failing SAM tests for the remaining CMS SRM endpoint. Eventually the problem was tracked down to a bug in the test. CMS fixed it and now we are back to our 'normal' level of SRM test failures.
- There was an issue with a network switch connecting some of the ECHO (CEPH) headnodes to the network on Friday (10th Feb). It was resolved when the switch's link was restarted on Tuesday morning.
Resolved Disk Server Issues
|
Current operational status and issues
|
- We are still seeing a rate of failures of the CMS SAM tests against the SRM. These are affecting our (CMS) availabilities but the level of failure have been reduced recently.
Ongoing Disk Server Issues
|
- GDSS663 (AtlasTape - D0T1) crashed on Saturday (18th Feb). Two faulty disks found and replaced. Expected back in service imminently.
Notable Changes made since the last meeting.
|
- The ECHO (CEPH) instance was upgraded yesterday (Tuesday 14th) to kraken.
Service
|
Scheduled?
|
Outage/At Risk
|
Start
|
End
|
Duration
|
Reason
|
Whole site
|
SCHEDULED
|
WARNING
|
01/03/2017 07:00
|
01/03/2017 11:00
|
4 hours
|
Warning on site during network intervention in preparation for IPv6.
|
Advanced warning for other interventions
|
The following items are being discussed and are still to be formally scheduled and announced.
|
Pending - but not yet formally announced:
- Merge AtlasScratchDisk into larger Atlas disk pool.
Listing by category:
- Castor:
- Update SRMs to new version, including updating to SL6. This will be done after the Castor 2.1.15 update.
- Networking:
- Enabling IPv6 onto production network.
- Databases
- Removal of "asmlib" layer on Oracle database nodes.
Entries in GOC DB starting since the last report.
|
Service
|
Scheduled?
|
Outage/At Risk
|
Start
|
End
|
Duration
|
Reason
|
lfc.gridpp.rl.ac.uk
|
SCHEDULED
|
WARNING
|
22/02/2017 08:45
|
22/02/2017 13:00
|
4 hours and 15 minutes
|
LFC Oracle backend security updates
|
All Castor and ECHO storage and Perfsonar.
|
SCHEDULED
|
WARNING
|
22/02/2017 07:00
|
22/02/2017 11:00
|
4 hours
|
Warning on Storage and Perfsonar during network intervention in preparation for IPv6.
|
Open GGUS Tickets (Snapshot during morning of meeting)
|
GGUS ID |
Level |
Urgency |
State |
Creation |
Last Update |
VO |
Subject
|
1267
|
Green
|
Very Urgent
|
In Progress
|
2017-02-22
|
2017-02-22
|
LHCb
|
File access problem at RAL
|
126718
|
Green
|
Urgent
|
In Progress
|
2017-02-21
|
2017-02-21
|
Atlas
|
UK RAL-LCG2-ECHO DATADISK: ~8k deletion error due to "Device or resource busy"
|
126532
|
Green
|
Urgent
|
In Progress
|
2017-02-09
|
2017-02-21
|
Atlas
|
RAL tape staging errors
|
126184
|
Green
|
Less Urgent
|
In Progress
|
2017-01-26
|
2017-02-07
|
Atlas
|
Request of inputs for new sites monitoring
|
124876
|
Red
|
Less Urgent
|
On Hold
|
2016-11-07
|
2017-01-01
|
OPS
|
[Rod Dashboard] Issue detected : hr.srce.GridFTP-Transfer-ops@gridftp.echo.stfc.ac.uk
|
117683
|
Red
|
Less Urgent
|
On Hold
|
2015-11-18
|
2017-02-10
|
|
CASTOR at RAL not publishing GLUE 2. Looking at it again now (Feb), progress made on back end. Need to update ticket.
|
Key: Atlas HC = Atlas HammerCloud (Queue ANALY_RAL_SL6, Template 845); Atlas HC ECHO = Atlas ECHO (Template 842);CMS HC = CMS HammerCloud
Day |
OPS |
Alice |
Atlas |
CMS |
LHCb |
Atlas HC |
Atlas HC ECHO |
CMS HC |
Comment
|
15/02/17 |
100 |
100 |
100 |
96 |
100 |
99 |
100 |
100 |
Timeouts on CMS SRM tests.
|
16/02/17 |
100 |
100 |
100 |
92 |
100 |
100 |
100 |
100 |
Timeouts on CMS SRM tests.
|
17/02/17 |
100 |
100 |
100 |
88 |
100 |
100 |
99 |
100 |
Timeouts on CMS SRM tests.
|
18/02/17 |
100 |
100 |
100 |
97 |
100 |
100 |
96 |
100 |
Timeouts on CMS SRM tests.
|
19/02/17 |
100 |
100 |
100 |
97 |
100 |
100 |
99 |
100 |
Timeouts on CMS SRM tests.
|
20/02/17 |
100 |
100 |
100 |
96 |
100 |
98 |
97 |
100 |
Timeouts on CMS SRM tests.
|
21/02/17 |
100 |
100 |
100 |
98 |
100 |
98 |
93 |
100 |
Timeouts on CMS SRM tests.
|