Difference between revisions of "Tier1 Operations Report 2015-11-11"

From GridPP Wiki
Jump to: navigation, search
(Created page with "==RAL Tier1 Operations Report for 11th November 2015== __NOTOC__ ====== ====== <!-- ************************************************************* -----> <!-- ***********Start...")
 
()
Line 121: Line 121:
 
|-style="background:#b7f1ce"
 
|-style="background:#b7f1ce"
 
! GGUS ID !! Level !! Urgency !! State !! Creation !! Last Update !! VO !! Subject
 
! GGUS ID !! Level !! Urgency !! State !! Creation !! Last Update !! VO !! Subject
|-
 
| 117277
 
| Green
 
| Urgent
 
| Waiting for Reply
 
| 2015-10-30
 
| 2015-11-03
 
| Atlas
 
| UK RAL-LCG2 staging error "bring-online timeout has been exceeded" (over 500 errors)
 
|-
 
| 117248
 
| Green
 
| Less Urgent
 
| In Progress
 
| 2015-10-28
 
| 2015-11-03
 
|
 
| Incorrect Certificate for SRM (new certs requested but not yet deployed)
 
 
|-
 
|-
 
| 116866
 
| 116866
Line 152: Line 134:
 
| Green
 
| Green
 
| Urgent
 
| Urgent
| In Progress
+
| On Hold
 
| 2015-10-12
 
| 2015-10-12
| 2015-10-26
+
| 2015-11-04
 
| CMS
 
| CMS
 
| T1_UK_RAL AAA opening and reading test failing again...
 
| T1_UK_RAL AAA opening and reading test failing again...

Revision as of 12:23, 11 November 2015

RAL Tier1 Operations Report for 11th November 2015

Review of Issues during the week 4th to 11th November 2015.
  • We again saw high load on the AtlasTape Castor instance - exacerbated by the failure of some disk servers in the cache in front of this tape area.
Resolved Disk Server Issues
  • gdss665 (AtlasTape - D0T1) failed on Sat (24th Oct). Following a disk replacement and updating the firmware in the disk controller the system was re-run through the acceptance testing for 5 days before being returned to service yesterday (3rd November).
  • gdss663 (AtlasTape - D0T1) failed on Sun (25th Oct). Following a disk and battery replacement and updating the firmware in the disk controller the system was re-run through the acceptance testing for 5 days before being returned to service yesterday (3rd November).
Current operational status and issues
  • The LHCb problem with a low but persistent rate of failure when copying the results of batch jobs to Castor. There is also a further problem that sometimes occurs when these (failed) writes are attempted to storage at other sites.
  • The intermittent, low-level, load-related packet loss seen over external connections is still being tracked. Likewise we have been working to understand some remaining low level of packet loss seen within a part of our Tier1 network.
  • Long-standing CMS issues. The two items that remain are CMS Xroot (AAA) redirection and file open times. Work is ongoing into the Xroot redirection with a new server having been added in recent weeks. File open times using Xroot remain slow but this is a less significant problem.
Ongoing Disk Server Issues
  • gdss707 (AtlasDataDisk - D1T0) has been out of production since Friday (16th Oct). The server was drained and is currently undergoing testing with fabric.
  • gdss664 (AtlasTape - D0T1) was removed from service on the 28th Oct. The system was having some problems running some network commands. These were resolved by a reboot. A failing disk was also replaced. The system has had the disk controller firmware updated and has been re-running the acceptance tests since yesterday (3rd Nov).
Notable Changes made since the last meeting.
  • None
Declared in the GOC DB
  • None
Advanced warning for other interventions
The following items are being discussed and are still to be formally scheduled and announced.
  • Upgrade of remaining Castor disk servers (those in tape-backed service classes) to SL6. This will be transparent to users.
  • Some detailed internal network re-configurations to enable the removal of the old 'core' switch from our network. This includes changing the way the UKLIGHT router connects into the Tier1 network.

Listing by category:

  • Databases:
    • Switch LFC/3D to new Database Infrastructure.
  • Castor:
    • Update SRMs to new version (includes updating to SL6).
    • Update disk servers to SL6 (ongoing)
    • Update to Castor version 2.1.15.
  • Networking:
    • Complete changes needed to remove the old core switch from the Tier1 network.
    • Make routing changes to allow the removal of the UKLight Router.
  • Fabric
    • Firmware updates on remaining EMC disk arrays (Castor, LFC)
Entries in GOC DB starting since the last report.
  • None
Open GGUS Tickets (Snapshot during morning of meeting)
GGUS ID Level Urgency State Creation Last Update VO Subject
116866 Green Less Urgent On Hold 2015-10-12 2015-10-19 SNO+ snoplus support at RAL-LCG2 (pilot role)
116864 Green Urgent On Hold 2015-10-12 2015-11-04 CMS T1_UK_RAL AAA opening and reading test failing again...
Availability Report

Key: Atlas HC = Atlas HammerCloud (Queue ANALY_RAL_SL6, Template 508); CMS HC = CMS HammerCloud

Day OPS Alice Atlas CMS LHCb Atlas HC CMS HC Comment
04/11/15 100 100 98 100 100 92 100 Single SRM test failure "could not open connection to srm-atlas.gridpp.rl.ac.uk:8443"
05/11/15 100 100 100 100 100 93 100
06/11/15 100 100 100 100 100 90 100
07/11/15 100 100 100 100 100 97 100
08/11/15 100 100 100 100 100 100 N/A
09/11/15 100 100 100 100 100 93 100
10/11/15 100 100 100 100 100 100 100