Dedicated model · Available as managed deployment
Meta's Llama 3.1 family — 8B, 70B and 405B instruction-tuned models with a 128k-token context and multilingual training — the most deployed open LLM family in production. Validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only — an OpenAI-compatible endpoint on hardware only you use, operated by AxForge in the EU.
Why AxForge
| The default open model | Llama 3.1 is what most teams standardise on: mature tooling, broad framework support and years of community fine-tunes and evaluations. |
|---|---|
| 8B and 70B on one machine | The 8B runs comfortably on a dedicated DGX Spark; the 70B fits the same machine in a quantized build. 405B is a multi-GPU deployment. |
| Licence handled | The Llama 3.1 community licence allows commercial use under its conditions; AxForge deploys under it and tells you exactly what applies to you. |
Specifications
| Model | Llama 3.1 — meta-llama |
|---|---|
| Modalities | Text |
| Sizes | 8.0B, 70.6B, 405.9B |
| Licence | Open, with conditions — llama3.1; AxForge deploys under it and tells you what applies |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Rental term | Hour, week, month or year |
| Hardware pricing | €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT |
| Managed service | Quoted per deployment |
| Region | Málaga, Spain (eu-es-1) |
Full details, benchmarks and FAQ on the Llama 3.1 page. Prices exclude VAT.
How it works
| 1 | Request deployment — describe your traffic, context needs and rental term. |
|---|---|
| 2 | You receive the configuration, hardware rental and managed-service price in writing before anything is billed. |
| 3 | AxForge deploys Llama 3.1 on a dedicated DGX Spark reserved for you. |
| 4 | Point your OpenAI SDK at your own endpoint with the model name you receive. |
| 5 | Adjust the term — hour, week, month or year — as your workload settles. |
Request deployment or sign in to start.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.
8B natively and 70B in a quantized build, both with the 128k context; 405B needs a multi-GPU system, which AxForge scopes with you.
Under the Llama 3.1 community licence, yes for almost every business — the licence carries conditions, and AxForge walks you through them before deployment.
AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.
Hardware by the hour, week, month or year; the managed service is quoted per deployment — both confirmed in writing before anything is billed.