[Jlab-scicomp-briefs] Status update: Unplanned Service Outage - SciComp/Farm

Brad Sawatzky brads at jlab.org
Wed Aug 5 18:30:29 EDT 2026


We have been making solid progress on bringing the SciComp filesystems
back online throughout the day.  The filesystems backing /work,
code.jlab.org CI/CD services, etc seem to have weathered the event
fairly well.  Restoration of the Lustre-backed filesystems /cache and
/volatile (ENP and LQCD) has been slowed by hardware issues and is
taking longer to recover.  This process will continue into Thursday.

Given how integral the lustre filesystems are to operations on both the
Farm and LQCD clusters, we are keeping those systems offline for another
day.  Please do not try to login to ifarm or the interactive QCD nodes;
even if you catch a window and log in you will be missing filesystems at
best.

Our goal is to start bringing the Farm and LQCD clusters online by EOB
on Thursday.  Globus, Rucio, and XrootD access will be restored at that
time as well.

Reminders For Hall Operations:
  - Hall systems that mounted central filesystems may require a force
    unmount and/or a reboot.
    - Note that the /lustre or /v mounts are not available, so you
      can't reach /cache or /volatile.  Services or shells may stall
      if they attempt to access those mounts.
  - 'jmirror' and 'jput' operations from the Hall clusters are
    working, but tape jobs may queue while we work though some issues.
    - Note that the 'autocopy' to /cache feature will not happen until
      /cache is restored, *but* the files will still make it to tape.
      The /mss interface to the tape system file inventory is
      operational.

We will provide further updates on Thursday.

Thanks again for bearing with us..

On behalf of Scientific Computing Operations,
-- Brad


More information about the Jlab-scicomp-briefs mailing list