Difference between revisions of "Tier1 Operations Report 2017-02-15"

Revision as of 23:36, 14 February 2017

Review of Issues during the week 8th to 15th February 2017.

There was also a problem with ALICE after the 'GEN' upgrade. ALICE require a special version of the xroot component for Castor. Checks that the xroot component would install under 2.1.15 had been made - but a newer version was needed. Once this had been provided there was a further ALICE specific configuration error that had to be tracked down. This caused a significant loss of availability for ALICE (failed between the 26th and 30th January).
Since the castor upgrade we have seen a couple of further problems:
- There has been a problem with the LHCb instance - we see a database resource (number of cursors) exhausted - and have had to restart the service to clear stuck transfers (on 1st Feb). A similar operation was carried out for Atlas (on 31st Jan).
- We have been failing tests for CMS xroot redirection. This appears to have started a couple of days after the CMS stager upgrade and is not yet understood.
- We are also failing CMS tests for an SRM endpoint defined in the GOC DB but not in production ("srm-cms-disk"). This should not have tests running against it and needs following up with CMS. Even though this test should not matter we would like to understand why it has stopped working after the Castor upgrade.

Resolved Disk Server Issues

GDSS674 (CMSTape - D0T1) reported problems on Friday evening, 10th Feb. There were no files on the server awaiting to go to tape. After being checked over it was returned to service on Monday (13th).

Current operational status and issues

We have been seeing a rate of failures of the CMS SAM tests against the SRM. These are affecting our (CMS) availabilities. A correction to the list of CMS 'services' being tests is helping with the resulting availability measure.

Ongoing Disk Server Issues

Notable Changes made since the last meeting.

Declared in the GOC DB

Service	Scheduled?	Outage/At Risk	Start	End	Duration	Reason
Whole site	SCHEDULED	WARNING	01/03/2017 07:00	01/03/2017 11:00	4 hours	Warning on site during network intervention in preparation for IPv6.
All Castor and ECHO storage and Perfsonar.	SCHEDULED	WARNING	22/02/2017 07:00	22/02/2017 11:00	4 hours	Warning on Storage and Perfsonar during network intervention in preparation for IPv6.

Advanced warning for other interventions

The following items are being discussed and are still to be formally scheduled and announced.

Pending - but not yet formally announced:

Listing by category:

Castor:
- Update SRMs to new version, including updating to SL6. This will be done after the Castor 2.1.15 update.
Networking:
- Enabling IPv6 onto production network.
Databases
- Removal of "asmlib" layer on Oracle database nodes.

Entries in GOC DB starting since the last report.

Service	Scheduled?	Outage/At Risk	Start	End	Duration	Reason
ECHO: gridftp.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, s3.echo.stfc.ac.uk, xrootd.echo.stfc.ac.uk	UNSCHEDULED	OUTAGE	13/02/2017 00:00	13/02/2017 13:45	13 hours and 45 minutes	Problem with switch causing Echo to stop being accessible.

Open GGUS Tickets (Snapshot during morning of meeting)


GGUS ID	Level	Urgency	State	Creation	Last Update	VO	Subject
126533	Green	Urgent	In Progress	2017-02-10	2017-02-14	Atlas	UK RAL-LCG2-ECHO transfer/staging/deletion failures with "Unable to connect to gridftp.echo.stfc.ac.uk"
126532	Green	Urgent	In Progress	2017-02-09	2017-02-10	Atlas	RAL tape staging errors
126184	Green	Less Urgent	In Progress	2017-01-26	2017-02-07	Atlas	Request of inputs for new sites monitoring
124876	Red	Less Urgent	On Hold	2016-11-07	2017-01-01	OPS	[Rod Dashboard] Issue detected : hr.srce.GridFTP-Transfer-ops@gridftp.echo.stfc.ac.uk
117683	Red	Less Urgent	On Hold	2015-11-18	2016-12-07		CASTOR at RAL not publishing GLUE 2. We looked at this as planned in December (report).

Availability Report

Key: Atlas HC = Atlas HammerCloud (Queue ANALY_RAL_SL6, Template 845); CMS HC = CMS HammerCloud


Day	OPS	Alice	Atlas	CMS	LHCb	Atlas HC	CMS HC	Comment
08/02/17	100	100	88	100	100	100	99	SRM test failures (timeouts)
09/02/17	100	100	59	100	100	99	100	SRM test failures (timeouts)
10/02/17	100	100	21	100	100	98	100	SRM test failures (timeouts)
11/02/17	100	100	100	100	100	100	100
12/02/17	100	100	100	100	100	100	100
13/02/17	100	100	100	100	100	100	100
14/02/17	100	100	100	100	100	100	100

Notes from Meeting.

@@ Line 35: / Line 35: @@
 | style="background-color: #b7f1ce; border-bottom: 1px solid silver; text-align: center; font-size: 1em; font-weight: bold; margin-top: 0; margin-bottom: 0; padding-top: 0.1em; padding-bottom: 0.1em;" | Current operational status and issues
 |}
-* There is a problem seen by LHCb of a low but persistent rate of failure when copying the results of batch jobs to Castor. There is also a further problem that sometimes occurs when these (failed) writes are attempted to storage at other sites.
 * We have been seeing a rate of failures of the CMS SAM tests against the SRM. These are affecting our (CMS) availabilities. A correction to the list of CMS 'services' being tests is helping with the resulting availability measure.
 <!-- ***********End Current operational status and issues*********** ----->