Hi,
We use option 1, but every 2 minutes (for jobs, but also everything else: users, accounts, nodes, reservations…). We also add data from other sources (CMDB, PuppetDB, ES), so we have a comprehensive data source containing everything we might want to know.
This database is then exposed through a REST API and provides two web interfaces: one for users and one for administrators. We basically have our own Slurm-web, but without any risk of overloading slurmdbd, and with a 2-minute delay that can be manually bypassed by reloading the data.
The Slurm database is purged every month, but we keep all records in our own database. (Don't forget to partition it!)
Thanks for the suggestion. I do like option 1 as it seems to me to get what you would want which a more detailed version of the job completion logs. I'm actually surprised there isn't a way that I can see to tune what fields you get in the job completion logs, as that would see to be an ideal method for doing this (almost what it was designed for).
-Paul Edmon-
Hi Paul,
Three approaches I've seen work, roughly from least to most effort:
1. Nightly export to your own warehouse. Once a day, run sacct --parsable2 --allusers --duplicates for the previous day (well inside your 1-week query limit) and write it to Parquet or Postgres somewhere offsite. It's one small query per day, so it barely touches the production dbd, and nothing in it ever gets purged. For history, load your existing archive files into a throwaway slurmdbd with sacctmgr archive load, export those once the same way, and you have one continuous record.
2. A jobcomp plugin. JobCompType=jobcomp/elasticsearch (or jobcomp/kafka on recent versions) streams each finished job record off the controller asynchronously, independent of slurmdbd and its purge schedule. The catch is that jobcomp records carry fewer fields than the full accounting tables, with less step-level detail, so check they cover what you need for billing.
3. Change-data-capture from the MySQL binlog. A CDC tool such as Debezium or Maxwell reads the binlog and applies inserts and updates to an offsite copy, and you drop the DELETE events the purge generates. It's the closest thing to "a replica that ignores purges" and doesn't slow the cluster, but it's more moving parts, and schema changes on Slurm upgrades need watching.
For raw data you can reprocess in new ways, option 1 is the one I'd start with. It's simple, it's plain files, and it survives upgrades.
Jim Jardine
On Fri, Sep 25, 2026 09:52 AM, Paul Edmon via slurm-users <slurm-users@lists.schedmd.com> wrote:
Currently we have a slurmdbd that is our main production, and thus under
heavy traffic with query restrictions (1 week at a time currently). We
also archive data that is older than 6 months so that the database stays
lean and mean. This serves the purpose well for general day to day but
limits us in terms of doing larger queries for the sake of summary
statistics and billing. What I would like to stand up is a second
slurmdbd that has no query restrictions and has all my historical data,
while also being kept abreast of the live information coming off the
cluster.
Straight replication (slurmdbd or mysql) won't suffice as when I do my
monthly purges, it would then propagate that purge. Likewise reimporting
the archived data into an external database would be missing the current
6 months. My question is is there a clever way to run a duplicate
slurmdbd that gets all the cluster traffic so it stays up to date, while
also not purging data? Also in such as way that I could say have the
duplicate slurmdbd offsite and not drag down cluster performance (as
really the cluster should lean on the active database that is more
proximate)? Looking at the documentation I can see an obvious method for
HA, and AccountingStorageExternalHost which looks like it might work but
may drag down performance due to needing to wait for DB operations to
complete there.
I'm curious if other sites have solved this. Summary programs like XDMod
are useful but having the raw data available to cross check and process
in new ways is what we are really looking for.
Thanks in advance.
-Paul Edmon-
--
slurm-users mailing list -- slurm-users@lists.schedmd.com
To unsubscribe send an email to slurm-users-leave@lists.schedmd.com