RAL Tier1 weekly operations castor 08/03/2018

From GridPP Wiki
Revision as of 14:18, 8 March 2018 by Rob Appleyard 7f7797b74a (Talk | contribs)

(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to: navigation, search

Draft agenda

1. Problems encountered this week

2. Upgrades/improvements made this week

3. What are we planning to do next week?

4. Long-term project updates (if not already covered)

  1. SL5 elimination from CIP and tape verification server
  2. CASTOR stress test improvement
  3. Generic CASTOR headnode setup
  4. Aquilonised headnodes

5. Special topics

6. Actions

7. Review Fabric tasks

  1.   Link

8. AoTechnicalB

9. Availability for next week

10. On-Call

11. AoOtherB

Operation problems

  * fdsdss34 out for a week and was not noticed >:(. 
     * Failed one evening, put in Nagios downtime, no action taken and ticket closed.
     * Now fixed.
  * Some tape drives got stuck in busy state after Facilities DB upgrade.
     * Chris will raise a ticket so it has been recorded.
  * Chris reports that a functional test that used to run on lcgcadm05 doesn't work anymore (probably because the machine has been turned off)
     * Should have been ported to castor-functional-test1 - maybe just needs repointing?
  * Workload increase on Neptune - ATLAS_STAGER - to be investigated.

Operation news

  * Oracle patching done on Facilities. Completed comfortably without the intervention window.
  * New filebeats for CASTOR daemons developed

Plans for next few weeks

Patching of the Neptune and Pluto DB and testing of switching over to R26 postponed until 27th March.

RA: Be in Japan and Taiwan GP: Fix-up on Aquilon SRM profiles:

  1. Move nscd feature to a sub-dir task - awaiting deployment
  2. Make castor/cron-jobs/srmbed-monitoring part of castor/daemons/srmbed feature task - awaiting deployment.

GP: Continue work on 'macroheadnodes' and write the change control. Deployment of new genTape disk servers.

Upgrade vcert to 2.1.16-13

Long-term projects

Headnode migration to Aquilon - Stager, scheduler, utility and nameserver configuration mainly complete. Stager, scheduler, utility tested seperately and all together on preprod. Will start combining the Stager, scheduler, utility features in one node.

'Macroheadnode' configuration testable :D

HA-proxyfication of the CASTOR SRMs: HA proxy is back and can be tested on preprod

Target: Combined headnodes running on SL7/Aquilon - implement CERN-style 'Macro' headnodes.

Draining of 4 x 13 generation disk servers from Atlas that will be deployed on genTape - waiting for Fabric.

Draining of 10% of the 14 generation disk servers

Actions

RA/BD: Run GFAL unit tests against CASTOR. Get them here: https://gitlab.cern.ch/dmc/gfal2/tree/develop/test/

RA to organise a meeting with the Fabric team to discuss outstanding issues with Data Services hardware

GP to talk to Alastair about draining of 10% of 14 gen disk servers

GP/RA to write a Nagios test to check for large number of requests that remain for a long time

Staffing

GP on call RA out from Friday until 26th March.