Ollama
Ollama server on port 11434 (OLLAMA_HOST=0.0.0.0), plus SSH. Pull any model with `ollama pull`.
| Detail | Value |
|---|---|
| Image | ghcr.io/fairgpu/ollama:latest |
| Mode | interactive |
| Ports | 22/ssh, 11434/http |
| Needs | 8 GB VRAM, 40 GB disk |
| Compatible machines online | 1 |
| Used | 0 times |
Ollama
The Ollama API listens on port 11434 (OLLAMA_HOST=0.0.0.0). Over SSH: ollama pull llama3.1:8b, then curl http://<relay>:<port>/api/generate -d '{"model":"llama3.1:8b","prompt":"hi"}' from your machine.
- VRAM: 8 GB for 7-8B Q4 models, 24 GB for 30B+ quantised, 48 GB+ for 70B.
- Set
OLLAMA_PULL(e.g.llama3.1:8b) to pull a model at start.
Before you stop
Everything you create lives inside the container on the host's disk and is removed when the rental ends. Push results out before stopping (scp, rsync, git, HF Hub, S3...). From inside the container you can run fairgpu extend 60 to add time, fairgpu mark "epoch 3 done" to annotate the timeline and fairgpu stop when you are done.
Step by step
1. Pull a model
Open the browser Terminal (or SSH) and run:
ollama pull llama3.1:8b # ~5 GB, fits 8 GB VRAM
ollama list
Models are stored in /workspace/ollama (OLLAMA_MODELS), so they are kept with a persistent workspace or checkpoint. Set OLLAMA_PULL=llama3.1:8b in the template env to pull at start.
2. Talk to it from your computer
The API port (11434) is public on your relay port:
curl http://<relay>:<port>/api/generate -d '{"model":"llama3.1:8b","prompt":"Explain attention in two sentences","stream":false}'
curl http://<relay>:<port>/v1/chat/completions -H 'content-type: application/json' -d '{"model":"llama3.1:8b","messages":[{"role":"user","content":"hi"}]}'
Point any OpenAI-compatible client at http://<relay>:<port>/v1 (any API key). Anyone with the URL can use it while the rental runs - stop the rental when done.
3. Pick a size for your VRAM
8 GB: 7-8B Q4 (llama3.1:8b, qwen2.5:7b); 24 GB: 30B+ quantised (qwen2.5:32b); 48 GB+: 70B (llama3.1:70b). ollama ps shows what is loaded and on which device.
4. Chat in the terminal
ollama run llama3.1:8b - type /bye to exit.
Save your results before you stop
- Everything outside
/workspacedisappears when the rental stops;/workspacesurvives on the same host (persistent workspace) and anywhere with a FairGPU checkpoint. - Save checkpoint now on the rental page (or
fairgpu snapshot save "label"inside the container) archives /workspace to FairGPU cloud storage; Download workspace gets it to your computer. scp -P <port> root@<relay>:/workspace/results ./copies files out over SSH;rsyncand SFTP (WinSCP/FileZilla) work the same way.- Stop the rental to stop billing. Set an idle auto-stop if you tend to forget, or buy the $0.99 finish alert (SMS + email).
Machines that can run it
| Machine | GPU | VRAM | Price | Status | Reliability | Country | Host | Cached |
|---|---|---|---|---|---|---|---|---|
| WhiteBob | NVIDIA GeForce RTX 5060 Ti | 16 GB | $0.10/hr | available now | 62 % | - | Aharon Sela | - |