xmrcloud
menu

GPU Pro — RTX A6000 — no-KYC offshore GPU servers (Iceland, Romania, Monero)

gpu-pro — Workstation-class GPU for ML training offshore.

xmrcloud-cli provision --plan=gpu-pro --region=<is|ro>

Order GPU Pro — RTX A6000

monthly $1099
annual -20% $10550
biennial -25% $19782
order gpu-pro

no-kyc crypto billing (xmr recommended; btc / ltn / ltc / eth / usdt accepted) — why-monero covers the rationale, payments the flow.

xmrcloud-cli spec --plan=gpu-pro

cpu 16 vCPU (EPYC)
ram 128 GB DDR5
storage 2 TB NVMe SSD
bandwidth 30 TB / month
port 10 Gbps
virtualization Bare-Metal
os Ubuntu 22.04 + CUDA 12, Custom
ip 1 × IPv4 + IPv6
ddos-shield 80 Gbps
uptime-sla 99.9%

notes

  • NVIDIA RTX A6000 (48 GB VRAM)
  • ECC memory

xmrcloud-cli describe --plan=gpu-pro

gpu-pro is an RTX A6000 (48 GB ECC GDDR6) on 16 EPYC cores, 128 GB DDR5, and 2 TB NVMe. The 48 GB VRAM is the practical floor for LoRA / QLoRA fine-tuning of 13B–34B models and for serving a 70B model quantised to 4-bit under vLLM — work that overflows gpu-lite's 24 GB 4090. Full-precision 70B training and HBM-bandwidth-bound inference still want gpu-beast's H100. ECC memory matters here: multi-hour training runs are where silent bit-flips actually surface. Iceland-hosted, Monero-billed.

xmrcloud-cli regions --plan=gpu-pro

region country ping flag
is Iceland (Reykjavik) ~38ms FRA --region=is
switzerland Switzerland (Zurich) ~?ms --region=switzerland
netherlands Netherlands (Amsterdam) ~?ms --region=netherlands

after you click order

xmrcloud-cli provision --plan=gpu-pro --region=is
[ok] reserving capacity in region=is
[ok] node allocated: gpu-pro-is-21
[ok] applying hardened-by-default profile (sshd, fail2ban, unattended-upgrades)
[ok] base image bootstrapped (Debian 12)
[ok] handoff key sealed → view via the console at /console
provisioned in 47s. ssh access via onion-auth or wireguard, your choice.

you receive the onion-auth key + initial sshd config in the same handoff. no email-shipped credentials. nothing is logged to the operator side.

cat /etc/xmrcloud/baseline.d/*

Every GPU Pro — RTX A6000 ships with the xmrcloud hardening baseline applied on the first boot — no opt-in flag, no add-on, no separate purchase. The baseline is the same across the catalog (vps / dedicated / gpu / tor / i2p / lokinet); category-specific extras are listed below the common section. Detailed per-control runbooks live in /docs; the cross-cutting overview is at /hardening.

  • KERNEL. KSPP-baseline sysctls applied (kernel.kptr_restrict=2, kernel.yama.ptrace_scope=1, kernel.unprivileged_bpf_disabled=1, vm.unprivileged_userfaultfd=0, net.ipv4.tcp_syncookies=1, +12 more), unprivileged user-namespace creation gated, kexec disabled at runtime. Full list and rationale: /docs/kernel-hardening-checklist.
  • SSHD. PasswordAuthentication no, ChallengeResponseAuthentication no, KbdInteractiveAuthentication no, PermitRootLogin prohibit-password, MaxAuthTries 3, Ed25519-only host keys (RSA host keys removed), legacy KEX / cipher / MAC families disabled. fail2ban preconfigured with the sshd-default ruleset. Runbook: /docs/harden-sshd; key migration: /docs/ssh-key-migration.
  • AUDIT. auditd enabled with the laurel-compatible default ruleset (auth, identity, network-config, time-change, mount, perm-mod). unattended-upgrades on for main/security only — feature releases stay operator-controlled. systemd-journald persistent storage with SystemMaxUse=512M.
  • NETWORK. Egress-default-permit (the box reaches the internet), ingress-default-deny (only sshd + the customer's declared services). Outbound port 25 (SMTP) closed by default; customers operating a real MTA request the lift via /contact with the reverse-DNS pointing to a domain they control. Dual-stack IPv4 + IPv6 (/64 routed). RIPE- allocated PI on Iceland and Romania.
  • MONITORING. node_exporter (Prometheus textfile exporter) listening on 127.0.0.1:9100 — the operator's monitoring scrapes via wireguard from the management VLAN, never from the public internet. Customers wanting their own metrics tap add a second exporter on a private interface.
  • INFERENCE STACK. CUDA 12.x toolkit, cuDNN, NCCL, NVIDIA driver matched to the installed accelerator, vLLM (latest stable), Ollama with model-registry mirror, llama.cpp (CUDA-compiled), PyTorch + transformers — preinstalled, container runtime ready.

the baseline is editorial-stable — when the operator changes a default, the change is logged in /notes with the rationale and the migration notes for boxes already in service. /hardening is the canonical pillar; /docs is the procedural manual.

faq -p gpu-pro

What can gpu-pro's 48 GB A6000 do that gpu-lite can't?

LoRA/QLoRA fine-tuning of 13B–34B models, and serving a 70B model quantised to 4-bit under vLLM — both overflow gpu-lite's 24 GB 4090. The A6000's 48 GB and ECC memory are the practical floor for multi-hour training. Full-precision 70B and HBM-bandwidth-bound inference still want gpu-beast's H100.

Does gpu-pro's ECC memory matter?

For multi-hour training runs, yes — ECC catches the silent bit-flips that corrupt a long run, which is exactly where consumer cards without ECC bite. For pure inference it matters less. That ECC (plus 48 GB) is the main reason to pick the A6000 over a second 4090-class card.

Can gpu-pro serve a 70B LLM?

Quantised to 4-bit, yes — a 70B model fits 48 GB with usable context under vLLM. Full-precision (FP16) 70B does not fit 48 GB and belongs on gpu-beast's 80 GB H100. For 13B–34B, gpu-pro serves full-precision comfortably.

What GPU is in each tier?

Tier-specific. The lite tier is RTX 4090 24GB (Ada Lovelace, 16384 CUDA cores, 82.58 TFLOPs FP16); pro is RTX A6000 48GB (Ampere, 10752 CUDA cores, 48 GB ECC); beast is H100 80GB SXM (Hopper, 14592 CUDA cores, 989.4 TFLOPs FP16) — the latter is procurement-sensitive. Spec is on the plan detail page; if the listed GPU is unavailable at order time the operator surfaces the substitute via /contact before charging.

What does "offshore GPU hosting" actually buy me vs cloud?

Two things. (1) Jurisdiction — the box runs in Iceland or Romania, neither of which is subject to US export-control orders that gate cloud GPU access for some workloads (open-weight LLMs from non-US authors, certain red-team / jailbreak-eval workloads, etc.). (2) Billing privacy — Monero / no-KYC vs cloud-card-on-file. The trade-off vs hyperscaler GPU is honestly documented at /playbook/ai-inference.

Is vLLM / Ollama / llama.cpp preinstalled?

Yes — the gpu-* tiers ship with vLLM (latest stable), Ollama (with the model registry mirror configured), llama.cpp (CUDA-compiled), CUDA 12.x toolkit, cuDNN, NCCL, PyTorch with CUDA support, and the HuggingFace transformers stack. Customers needing other inference engines (TensorRT-LLM, mlc-llm, exllama2) install via the package manager — the box is a normal Linux machine with NVIDIA driver + container runtime support.

Can I serve LLM endpoints publicly?

Yes — the AUP (/legal/aup) does not restrict serving open-weight LLM inference endpoints. Customers running a public chat / API surface should configure their own rate-limits and authentication; the operator does not provide a hosted gateway. Per /docs (Caddy-fronted vLLM is the common pattern), the xmrcloud hardening defaults already cover the OS layer.

Where is the GPU server hosted?

Iceland (Reykjavik, RIPE) — the GPU catalog is Iceland-only because the Romanian racks do not have the GPU-density power / cooling provisioned. Iceland's hydroelectric + geothermal power is the operator's preference for GPU workloads on cost and emissions footprint; jurisdictional posture is at /location/is.

Do I need to pay in Monero?

No. XMR is recommended; OxaPay accepts BTC, Lightning, LTC, ETH, and USDT. GPU-tier orders settle the same way VPS orders do — per-order Monero subaddress on XMR (MRL-0006), straight invoice on the transparent rails. No card, no fiat. The /why-monero rationale applies identically.

order gpu-pro

no-kyc crypto billing (xmr recommended; btc / ltn / ltc / eth / usdt accepted) — why-monero covers the rationale, payments the flow.

ls /guide