Chris, that's a fair point, and I overstated it. A job that checkpoints, runs to its limit and resubmits has done useful work, so counting all of its energy as waste is wrong. The tool reports energy "at stake" rather than energy saved partly for this reason, but saying the job "returns nothing" went too far.

I think the two cases can be told apart from accounting data alone. A checkpoint-restart chain usually looks like the same user resubmitting a near-identical job shortly after the timeout, and a genuine failure usually doesn't. I'll add that split, so the report shows TIMEOUT energy both with and without likely restart chains instead of a single number.

If you have a rough idea of what fraction of timeouts at your site are intentional, that would help me calibrate it.

Thanks for the correction.

Jim

On Mon, 21 Sept 2026 at 12:03, Christopher Samuel via slurm-users <slurm-users@lists.schedmd.com> wrote:
On 9/21/26 12:07 pm, James Jardine via slurm-users wrote:
>
> Jobs that ended in TIMEOUT were about 6.6% of jobs but roughly 45% of total
> job energy. That follows once you say it out loud: a timed-out job runs its
> full wall-clock request by definition, and then returns nothing.

That's not necessarily true, it ran to the timelimit true, but that
might be the users intention, they could be doing regular checkpoints
and then restarting from them to carry on their simulation/calculation.
We see that a lot.

So I would caution against drawing that sort of conclusion without
checking with the users running those jobs.

All the best,
Chris

--
slurm-users mailing list -- slurm-users@lists.schedmd.com
To unsubscribe send an email to slurm-users-leave@lists.schedmd.com