Troubleshooting

Connection refused, key rejected, host key changed, stuck in STARTING, out of disk, slow pulls, terminal not available, and other things that go wrong.

Connecting

ssh: connect to host … port …: Connection refused or a timeout

  1. Is the rental RUNNING and does the SSH port show a green dot in the Connect dialog? Ports come up a few seconds after RUNNING.
  2. Copy the command again: the port changes with every rental.
  3. Some networks block outbound ports in the 30000 range (corporate Wi-Fi, some hotels). Try a phone hotspot. The browser terminal uses the normal HTTPS port and works almost everywhere.

Permission denied (publickey)

Your client is offering a key that is not in the pod.

  • Which keys were injected? The SSH tab shows the key name; the rent wizard lets you pick.
  • Force the right key: ssh -i ~/.ssh/id_ed25519 -p <port> root@<relay> (Windows: -i $env:USERPROFILE\.ssh\id_ed25519).
  • Add the key you are using to the running pod: SSH tab → Add a key to this pod, or from the browser terminal append it to /root/.ssh/authorized_keys.
  • Pasted the private key by mistake? Generate a new pair; the old one is compromised.

WARNING: REMOTE HOST IDENTIFICATION HAS CHANGED!

Expected: the relay host name is the same for every rental but each container has its own host key. Compare the fingerprint in the SSH tab, then remove the stale entry:

Windows

ssh-keygen -R "[relay.fairgpu.io]:30123"

macOS

ssh-keygen -R "[relay.fairgpu.io]:30123"

Linux

ssh-keygen -R "[relay.fairgpu.io]:30123"

The config block from the VS Code tab already sets StrictHostKeyChecking accept-new and a throwaway known-hosts file for that alias.

This host runs an older agent; use SSH

The browser terminal needs a host agent that supports it; this host has not updated yet. Use the SSH tab (add a key to the pod if you have none), or pick a host with a newer agent next time.

Jupyter asks for a token

Use the full link from the Connect dialog (the token is inside), or run jupyter server list in the terminal to print it.

Web UI port never turns green

The service inside the container may bind to 127.0.0.1 instead of 0.0.0.0. Check its startup flags (--listen, --host 0.0.0.0). Until the server's HTTPS proxy is on, open HTTP ports through an SSH tunnel: ssh -L 8188:localhost:8188 -p <port> root@<relay>.

Starting

Stuck in STARTING

The host is pulling your image. Custom images of 10+ GB take minutes on a home connection; the timeline shows layer progress. You are billed from RUNNING, not before. Prefer machines with the ⚡ ready in seconds chip for the template, or lighter images.

FAILED with an error

Usually a wrong image name or tag, a private image, or a command that exits immediately. The host's error is shown on the page; you are not charged for failed starts. For jobs, check the stderr tab.

Setup script failed

The script's output is a system run named setup in the log viewer. Interactive rentals stay usable so you can fix it in the terminal; jobs fail. Common causes: apt-get install without apt-get update, a missing -y, network-dependent downloads. Make the script idempotent and set -euo pipefail.

Restore failed

The checkpoint could not be downloaded or extracted into the new workspace (too large for the host's disk, or a corrupt upload). Try a machine with more free disk or restore a different checkpoint; the rental is not charged for the failed start.

While running

No space left on device

Container scratch space is limited. Move caches and outputs under /workspace, pick a larger container disk in the wizard on hosts that offer it, grow the workspace under Workspaces, or delete temporary files (du -sh /workspace/* | sort -h). Workspace over quota stops the rental after 15 minutes unless usage drops.

GPU not visible (torch.cuda.is_available() is False)

Run nvidia-smi first. If it shows the card, the Python stack does not match the driver's CUDA version: use an official template or a pytorch/pytorch:*-cuda12.x image. If nvidia-smi fails inside the container, stop and rent another machine; report it in the review.

Slow uploads from my computer

Your upload speed is the limit. Download public data inside the container instead (wget, aria2c, huggingface-cli), and upload only what is unique to you. For large results, checkpoint to FairGPU storage and download later rather than streaming through a slow link while billing runs.

The rental stopped by itself

Check the stop reason on the page: auto-stopped while idle (your idle limit; change it in the sidebar), maximum duration reached, wallet balance ran out, host needed the PC back or host went offline (see Billing and refunds for what you got back, and Resume / Resume on another machine).

After

My files are gone

Only /workspace persists, and only if you kept a workspace (same host) or a checkpoint. Without either, the container is destroyed at stop. Next time: tick Keep a workspace, set a checkpoint interval, or push results out during the run (Move files).

Wallet not credited after a top-up

Wait a minute and refresh; Stripe's webhook applies the credit. Still missing: contact support with the receipt.

Hosting problems

See the troubleshooting section of Host: earn with your PC: agent offline, Docker cannot see the GPU, machine shows PAUSED.

All guides