vLLM (OpenAI-compatible)

vLLM OpenAI-compatible server on port 8000; set MODEL to a Hugging Face model id.

DetailValue
Imageghcr.io/fairgpu/vllm-openai:latest
Modeinteractive
Ports22/ssh, 8000/http
Needs16 GB VRAM, 60 GB disk
Compatible machines online0
Used0 times

vLLM (OpenAI-compatible server)

Starts vllm serve $MODEL on port 8000 (/v1/chat/completions, /v1/models, /health). The model is downloaded at start, so the first minutes are spent pulling weights (watch the logs).

  • MODEL: Hugging Face id (gated models need HF_TOKEN).
  • VLLM_ARGS: extra flags (--max-model-len, --quantization awq, --tensor-parallel-size 2...).
  • VRAM: 16 GB for 7B fp16, 24 GB for 7B with long context or 13B AWQ, 48 GB+ for 30B+.

Before you stop

Everything you create lives inside the container on the host's disk and is removed when the rental ends. Push results out before stopping (scp, rsync, git, HF Hub, S3...). From inside the container you can run fairgpu extend 60 to add time, fairgpu mark "epoch 3 done" to annotate the timeline and fairgpu stop when you are done.

Step by step

1. Pick the model before you book

Set MODEL to a Hugging Face id (default Qwen/Qwen2.5-7B-Instruct). Gated models (Llama, Gemma) need HF_TOKEN - store it once under Secrets and reference it as {{secret:HF_TOKEN}}.

2. Wait for the download

The first minutes download the weights: watch Logs on the rental page; the vLLM API port turns green when /health answers. 7B fp16 needs ~16 GB VRAM; add --quantization awq to VLLM_ARGS for AWQ repos.

3. Use the OpenAI-compatible endpoint

export OPENAI_BASE_URL=http://<relay>:<port>/v1
export OPENAI_API_KEY=anything
curl $OPENAI_BASE_URL/models
curl $OPENAI_BASE_URL/chat/completions -H 'content-type: application/json' -d '{"model":"Qwen/Qwen2.5-7B-Instruct","messages":[{"role":"user","content":"hello"}]}'

Python: from openai import OpenAI; client = OpenAI(base_url=..., api_key="x"). The endpoint is open to anyone who knows the relay port while the rental runs.

4. Tune

VLLM_ARGS: --max-model-len 8192 (longer context = more VRAM), --tensor-parallel-size 2 on 2-GPU machines, --gpu-memory-utilization 0.9. Restart by stopping/starting the rental or pkill -f vllm from Commands.

5. Keep the weights

Set HF_HOME=/workspace/hf in env so the download lands in the workspace and is kept (persistent workspace or checkpoint).

Save your results before you stop

  • Everything outside /workspace disappears when the rental stops; /workspace survives on the same host (persistent workspace) and anywhere with a FairGPU checkpoint.
  • Save checkpoint now on the rental page (or fairgpu snapshot save "label" inside the container) archives /workspace to FairGPU cloud storage; Download workspace gets it to your computer.
  • scp -P <port> root@<relay>:/workspace/results ./ copies files out over SSH; rsync and SFTP (WinSCP/FileZilla) work the same way.
  • Stop the rental to stop billing. Set an idle auto-stop if you tend to forget, or buy the $0.99 finish alert (SMS + email).

Pick a machine for vLLM (OpenAI-compatible)