Axolotl fine-tune (job)

Runs an Axolotl training config as a job and exits when done.

DetailValue
Imageaxolotlai/axolotl:main-latest
Modebatch job
Ports-
Needs24 GB VRAM, 80 GB disk
Compatible machines online0
Used0 times

Axolotl fine-tune (JOB)

Replace the command with your own (clone a repo with the config, then train). The job's stdout/stderr are in the rental logs; push adapters/checkpoints to the Hub or S3 inside the command because the container is removed when the command exits.

  • VRAM: 24 GB for 7-8B QLoRA, 48 GB+ for full fine-tunes.

Before you stop

Everything you create lives inside the container on the host's disk and is removed when the rental ends. Push results out before stopping (scp, rsync, git, HF Hub, S3...). From inside the container you can run fairgpu extend 60 to add time, fairgpu mark "epoch 3 done" to annotate the timeline and fairgpu stop when you are done.

Step by step

1. Prepare the config

The job runs accelerate launch -m axolotl.cli.train /workspace/config.yml. Put the config in place before the command, e.g. make the command:

git clone https://github.com/you/finetune-cfg /workspace/cfg && cp /workspace/cfg/qlora.yml /workspace/config.yml && accelerate launch -m axolotl.cli.train /workspace/config.yml

Set output_dir: /workspace/out in the config so adapters land in the workspace. HF_TOKEN / WANDB_API_KEY come from Secrets.

2. Watch it

Logs streams stdout/stderr; Metrics shows GPU memory. Set a finish alert or rely on the job email when the command exits.

3. Push results inside the command

The container is removed when the command exits (unless you turn on keep alive). Append to the command:

... && huggingface-cli upload you/my-adapter /workspace/out --repo-type model

or keep a persistent workspace / checkpoint and download /workspace/out afterwards.

4. Re-run with a different config

With keep alive on, run fairgpu run "accelerate launch -m axolotl.cli.train /workspace/config2.yml" (or the Commands panel) without starting a new rental; set an idle auto-stop so the container does not sit idle.

5. VRAM

24 GB: 7-8B QLoRA; 48 GB+: full fine-tunes or 13B LoRA; use gradient_checkpointing: true and micro_batch_size: 1 when memory is tight.

Save your results before you stop

  • Everything outside /workspace disappears when the rental stops; /workspace survives on the same host (persistent workspace) and anywhere with a FairGPU checkpoint.
  • Save checkpoint now on the rental page (or fairgpu snapshot save "label" inside the container) archives /workspace to FairGPU cloud storage; Download workspace gets it to your computer.
  • scp -P <port> root@<relay>:/workspace/results ./ copies files out over SSH; rsync and SFTP (WinSCP/FileZilla) work the same way.
  • Stop the rental to stop billing. Set an idle auto-stop if you tend to forget, or buy the $0.99 finish alert (SMS + email).

Pick a machine for Axolotl fine-tune (job)