RAL Tier1 weekly operations castor 23/03/2018

From GridPP Wiki
Jump to: navigation, search

Draft agenda

1. Problems encountered this week

2. Upgrades/improvements made this week

3. What are we planning to do next week?

4. Long-term project updates (if not already covered)

  1. SL5 elimination from CIP and tape verification server
  2. CASTOR stress test improvement
  3. Generic CASTOR headnode setup
  4. Aquilonised headnodes

5. Special topics

6. Actions

7. Review Fabric tasks

  1.   Link

8. AoTechnicalB

9. Availability for next week

10. On-Call

11. AoOtherB

Operation problems

  * Serious firewall misconfiguration resulted in the breaking of external xrootd access to RAL
  * Chris reports that a functional test that used to run on lcgcadm05 doesn't work anymore (probably because the machine has been turned off)
     * Should have been ported to castor-functional-test1 - maybe just needs repointing?

Operation news

  * Aquilon SRM profiles fixed up and deployed to prod
  * lcgcts12 (preprod) upgrade to SL7 properly attached to the preprod instance

Plans for next few weeks

Patching of the Neptune and Pluto DB and testing of switching over to R26 postponed until 27th March.

GP: Continue work on 'macroheadnodes' and write the change control. Deployment of new genTape disk servers.

Long-term projects

Headnode migration to Aquilon - Stager, scheduler, utility and nameserver configuration mainly complete. Stager, scheduler, utility tested seperately and all together on preprod. Combined Stager, scheduler, utility and SRM features in one node which passes functional tests

'Macroheadnode' configuration testable :D

HA-proxyfication of the CASTOR SRMs: HA proxy is back and can be tested on preprod

Target: Combined headnodes running on SL7/Aquilon - implement CERN-style 'Macro' headnodes.

Draining of 4 x 13 generation disk servers from Atlas that will be deployed on genTape - draining complete, waiting for Fabric.

Draining of 10% of the 14 generation disk servers

Actions

RA/BD: Run GFAL unit tests against CASTOR. Get them here: https://gitlab.cern.ch/dmc/gfal2/tree/develop/test/

RA to organise a meeting with the Fabric team to discuss outstanding issues with Data Services hardware

GP/RA to write a Nagios test to check for large number of requests that remain for a long time

Staffing

RA on call