Dedicated model · Available as managed deployment
A mixture-of-experts Qwen3 with a 65,536-token context window, validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. An OpenAI-compatible endpoint on your own machine, operated by AxForge in Málaga, Spain.
Why AxForge
| Efficient volume workloads | A mixture-of-experts LLM from the Qwen3 family with a 65,536-token context window. AxForge publishes only numbers it measures itself and has not benchmarked this model on its nodes yet — the newer Qwen3.6 35B A3B carries a community figure on its page. |
|---|---|
| Your machine, your endpoint | A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) reserved for you, running qwen3-30b-a3b behind an OpenAI-compatible /v1 that serves only your traffic. |
| EU-hosted, zero prompt retention | Hosted in Málaga, Spain (eu-es-1). Prompts and completions are processed in memory — not logged, not retained, never used to train — the same policy as the serverless API. |
Specifications
| Model | Qwen3 30B A3B — mixture-of-experts, Qwen3 family |
|---|---|
| Served model name | qwen3-30b-a3b |
| Architecture | MoE |
| Context window | 65,536 tokens |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Rental term | Hour, week, month or year |
| Hardware pricing | €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT |
| Managed service | Quoted per deployment |
| Region | Málaga, Spain (eu-es-1) |
Full details, benchmarks and FAQ on the Qwen3 30B A3B page. Prices exclude VAT.
How it works
| 1 | Request deployment — describe your traffic, context needs and rental term. |
|---|---|
| 2 | You receive the configuration, hardware rental and managed-service price in writing before anything is billed. |
| 3 | AxForge deploys Qwen3 30B A3B on a dedicated DGX Spark reserved for you. |
| 4 | Point your OpenAI SDK at your own endpoint with model qwen3-30b-a3b. |
| 5 | Adjust the term — hour, week, month or year — as your workload settles. |
Request deployment or sign in to start.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.
A mixture-of-experts LLM from the Qwen3 family with a 65,536-token context window.
AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card — the newer Qwen3.6 35B A3B has a community figure on its page.
A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented by the hour, week, month or year, running Qwen3 30B A3B behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.
Two parts: the DGX Spark hardware rental — by the hour, week, month or year, with longer terms earning the lower rate — and the managed service, quoted per deployment. Both are confirmed in writing before anything is billed.