Run a job and get notified

Batch work without babysitting: job mode, keep-alive containers, the Commands panel, idle auto-stop, the finish alert, and getting results out.

Job mode runs one command in a Docker image and streams the logs to your rental page. It is the right mode for training runs, renders, batch inference and anything you would otherwise start in tmux and walk away from.

Start a job

  1. Marketplace → pick a machine → Rent → mode Job.
  2. Image: a template (for example Axolotl fine-tune) or any public image such as pytorch/pytorch:2.4.1-cuda12.4-cudnn9-runtime.
  3. Command: runs with sh -c in /workspace. The Recent menu recalls commands you used before; ↑ cycles through them.
   bash -c "git clone https://github.com/you/repo && cd repo && pip install -r requirements.txt && python train.py --epochs 3"
  1. Env: HF_TOKEN, WANDB_API_KEY… Store them once as secrets and insert {{secret:HF_TOKEN}}; the value is resolved on the host at start and never shown again.
  2. Keep the container running after the command finishes (on by default): the rental stays up so you can inspect outputs, run another command or open the terminal. Off = classic batch: the rental stops the second the command exits.
  3. Maximum duration is your cost cap. Auto-stop when idle stops a keep-alive job that sits with no command running.

While it runs

  • Logs stream live; stdout / stderr / System tabs, search, download as .txt.
  • Commands panel (keep-alive jobs and every interactive rental): type a command, Enter. Each run gets its own log slice, exit code and duration; click a run to filter the log viewer to it. One command at a time; Kill sends SIGTERM then SIGKILL.
  • Markers: fairgpu mark "epoch 3 done" from inside the container, or Add marker on the timeline.
  • Live metrics: GPU util and memory, CPU, RAM, disk, network, every 5 s.
  • From inside the container the fairgpu helper (official images) can fairgpu stop, fairgpu extend 30, fairgpu run "python eval.py", fairgpu snapshot save "before eval".

Get notified

  • Email on start, stop, failure and idle warnings is free and on by default (prefs under Account → Notifications).
  • Finish alert (paid add-on, price shown in the wizard): one SMS + email the moment the command exits, the machine goes idle, or the rental stops, with the cost so far and Stop / Extend / Open buttons. Needs a verified phone number. Arm it in the wizard or later from the rental page sidebar. It is refunded automatically when the host is at fault.
  • COMMAND_FINISHED emails for each Commands-panel run can be switched on under Notifications → rentals.

Get results out

The container is deleted when the rental ends, so plan the exit:

  • Write outputs under /workspace and tick Keep a workspace (same host) or set a checkpoint interval so a copy lands in FairGPU storage every 15–60 minutes; both survive the stop. See Checkpoints and resuming.
  • Push from inside the job: Hugging Face Hub (huggingface-cli upload), S3 / R2 / B2 (aws s3 sync, rclone), W&B artifacts, or a cloud sync target with push on stop.
  • Download afterwards: Progress safety → Download workspace, or scp / rsync while the container is still up (Move files).

Interruptions

If the host's PC goes away mid-job you are refunded according to their policy (full refund on a Committed host or on any interruption without notice; setup time + the last 10 minutes on a reclaim with notice). Your workspace is kept for 24 hours and the rental shows Resume; with a checkpoint you can also Resume on another machine. Details: Billing and refunds.

Exit codes and stop reasons

Shown asMeaning
exit 0, job finishedThe command completed.
exit NThe command failed with code N; check the stderr tab.
maximum duration reachedYour cap stopped it; extend next time or raise the cap.
wallet balance ran outTop up; the rental page shows affordable minutes before you start.
auto-stopped while idleNo command and no activity for your idle limit.
host needed the PC back / host went offlineInterruption; see refunds above.

All guides