vLLM (OpenAI-compatible)
vLLM OpenAI-compatible server on port 8000; set MODEL to a Hugging Face model id.
| Detail | Value |
|---|---|
| Image | ghcr.io/fairgpu/vllm-openai:latest |
| Mode | interactive |
| Ports | 22/ssh, 8000/http |
| Needs | 16 GB VRAM, 60 GB disk |
| Compatible machines online | 0 |
| Used | 0 times |
vLLM (OpenAI-compatible server)
Starts vllm serve $MODEL on port 8000 (/v1/chat/completions, /v1/models, /health). The model is downloaded at start, so the first minutes are spent pulling weights (watch the logs).
MODEL: Hugging Face id (gated models needHF_TOKEN).VLLM_ARGS: extra flags (--max-model-len,--quantization awq,--tensor-parallel-size 2...).- VRAM: 16 GB for 7B fp16, 24 GB for 7B with long context or 13B AWQ, 48 GB+ for 30B+.
Before you stop
Everything you create lives inside the container on the host's disk and is removed when the rental ends. Push results out before stopping (scp, rsync, git, HF Hub, S3...). From inside the container you can run fairgpu extend 60 to add time, fairgpu mark "epoch 3 done" to annotate the timeline and fairgpu stop when you are done.
Step by step
1. Pick the model before you book
Set MODEL to a Hugging Face id (default Qwen/Qwen2.5-7B-Instruct). Gated models (Llama, Gemma) need HF_TOKEN - store it once under Secrets and reference it as {{secret:HF_TOKEN}}.
2. Wait for the download
The first minutes download the weights: watch Logs on the rental page; the vLLM API port turns green when /health answers. 7B fp16 needs ~16 GB VRAM; add --quantization awq to VLLM_ARGS for AWQ repos.
3. Use the OpenAI-compatible endpoint
export OPENAI_BASE_URL=http://<relay>:<port>/v1
export OPENAI_API_KEY=anything
curl $OPENAI_BASE_URL/models
curl $OPENAI_BASE_URL/chat/completions -H 'content-type: application/json' -d '{"model":"Qwen/Qwen2.5-7B-Instruct","messages":[{"role":"user","content":"hello"}]}'
Python: from openai import OpenAI; client = OpenAI(base_url=..., api_key="x"). The endpoint is open to anyone who knows the relay port while the rental runs.
4. Tune
VLLM_ARGS: --max-model-len 8192 (longer context = more VRAM), --tensor-parallel-size 2 on 2-GPU machines, --gpu-memory-utilization 0.9. Restart by stopping/starting the rental or pkill -f vllm from Commands.
5. Keep the weights
Set HF_HOME=/workspace/hf in env so the download lands in the workspace and is kept (persistent workspace or checkpoint).
Save your results before you stop
- Everything outside
/workspacedisappears when the rental stops;/workspacesurvives on the same host (persistent workspace) and anywhere with a FairGPU checkpoint. - Save checkpoint now on the rental page (or
fairgpu snapshot save "label"inside the container) archives /workspace to FairGPU cloud storage; Download workspace gets it to your computer. scp -P <port> root@<relay>:/workspace/results ./copies files out over SSH;rsyncand SFTP (WinSCP/FileZilla) work the same way.- Stop the rental to stop billing. Set an idle auto-stop if you tend to forget, or buy the $0.99 finish alert (SMS + email).