Difference between revisions of "RAL Tier1 weekly operations castor 16/03/2018"

From GridPP Wiki
Jump to: navigation, search
(Operation news)
(Operation problems)
Line 31: Line 31:
 
== Operation problems ==
 
== Operation problems ==
  
   * vdqmd stopped running on lcgcnsd01 not sending content vo vmdqmd.log - to be investigated
+
    
 
   * Chris reports that a functional test that used to run on lcgcadm05 doesn't work anymore (probably because the machine has been turned off)
 
   * Chris reports that a functional test that used to run on lcgcadm05 doesn't work anymore (probably because the machine has been turned off)
 
       * Should have been ported to castor-functional-test1 - maybe just needs repointing?
 
       * Should have been ported to castor-functional-test1 - maybe just needs repointing?
  
 
   * Workload increase on Neptune - ATLAS_STAGER - to be investigated.
 
   * Workload increase on Neptune - ATLAS_STAGER - to be investigated.
 +
  * vdqmd stopped running on lcgcnsd01 not sending content vo vmdqmd.log - to be investigated
 +
  * Alice SAM tests failed (probably because of aliceDisk is full)
  
 
== Operation news ==
 
== Operation news ==

Revision as of 14:55, 16 March 2018

Draft agenda

1. Problems encountered this week

2. Upgrades/improvements made this week

3. What are we planning to do next week?

4. Long-term project updates (if not already covered)

  1. SL5 elimination from CIP and tape verification server
  2. CASTOR stress test improvement
  3. Generic CASTOR headnode setup
  4. Aquilonised headnodes

5. Special topics

6. Actions

7. Review Fabric tasks

  1.   Link

8. AoTechnicalB

9. Availability for next week

10. On-Call

11. AoOtherB

Operation problems

  * Chris reports that a functional test that used to run on lcgcadm05 doesn't work anymore (probably because the machine has been turned off)
     * Should have been ported to castor-functional-test1 - maybe just needs repointing?
  * Workload increase on Neptune - ATLAS_STAGER - to be investigated.
  * vdqmd stopped running on lcgcnsd01 not sending content vo vmdqmd.log - to be investigated
  * Alice SAM tests failed (probably because of aliceDisk is full)

Operation news

  * vCert2 DB schemas upgraded to 2.1.16
  * gdss644 deployed into preprodTape
  * lcgcts02 (vCert) was upgraded to SL7

Plans for next few weeks

Patching of the Neptune and Pluto DB and testing of switching over to R26 postponed until 27th March.

GP: Fix-up on Aquilon SRM profiles:

  1. Move nscd feature to a sub-dir task - awaiting deployment
  2. Make castor/cron-jobs/srmbed-monitoring part of castor/daemons/srmbed feature task - awaiting deployment.

GP: Continue work on 'macroheadnodes' and write the change control. Deployment of new genTape disk servers.

Upgrade vcert to 2.1.16-13

Long-term projects

Headnode migration to Aquilon - Stager, scheduler, utility and nameserver configuration mainly complete. Stager, scheduler, utility tested seperately and all together on preprod. Will start combining the Stager, scheduler, utility features in one node.

'Macroheadnode' configuration testable :D

HA-proxyfication of the CASTOR SRMs: HA proxy is back and can be tested on preprod

Target: Combined headnodes running on SL7/Aquilon - implement CERN-style 'Macro' headnodes.

Draining of 4 x 13 generation disk servers from Atlas that will be deployed on genTape - draining complete, waiting for Fabric.

Draining of 10% of the 14 generation disk servers

Actions

RA/BD: Run GFAL unit tests against CASTOR. Get them here: https://gitlab.cern.ch/dmc/gfal2/tree/develop/test/

RA to organise a meeting with the Fabric team to discuss outstanding issues with Data Services hardware

GP to talk to Alastair about draining of 10% of 14 gen disk servers

GP/RA to write a Nagios test to check for large number of requests that remain for a long time

Staffing

GP on call RA out from Friday until 26th March.