Dedicated model · Available as managed deployment

Request a Qwen3 30B A3B deployment on your own DGX Spark

A mixture-of-experts Qwen3 with a 65,536-token context window, validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. An OpenAI-compatible endpoint on your own machine, operated by AxForge in Málaga, Spain.

eu-es-1 · Málaga Available as managed deployment 65,536 ctx · MoE Quoted per deployment qwen3-30b-a3b.axforge.ai
Request deploymentTalk to an engineerSign in €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT Hardware rental plus a managed service quoted per deployment — both confirmed in writing before anything is billed.

Why AxForge

Why Qwen3 30B A3B as a managed deployment

Efficient volume workloadsA mixture-of-experts LLM from the Qwen3 family with a 65,536-token context window. AxForge publishes only numbers it measures itself and has not benchmarked this model on its nodes yet — the newer Qwen3.6 35B A3B carries a community figure on its page.
Your machine, your endpointA dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) reserved for you, running qwen3-30b-a3b behind an OpenAI-compatible /v1 that serves only your traffic.
EU-hosted, zero prompt retentionHosted in Málaga, Spain (eu-es-1). Prompts and completions are processed in memory — not logged, not retained, never used to train — the same policy as the serverless API.

Specifications

What you get

ModelQwen3 30B A3B — mixture-of-experts, Qwen3 family
Served model nameqwen3-30b-a3b
ArchitectureMoE
Context window65,536 tokens
HardwareNVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge
Rental termHour, week, month or year
Hardware pricing€0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT
Managed serviceQuoted per deployment
RegionMálaga, Spain (eu-es-1)

Full details, benchmarks and FAQ on the Qwen3 30B A3B page. Prices exclude VAT.

How it works

From sign-in to running

1Request deployment — describe your traffic, context needs and rental term.
2You receive the configuration, hardware rental and managed-service price in writing before anything is billed.
3AxForge deploys Qwen3 30B A3B on a dedicated DGX Spark reserved for you.
4Point your OpenAI SDK at your own endpoint with model qwen3-30b-a3b.
5Adjust the term — hour, week, month or year — as your workload settles.

Request deployment or sign in to start.

FAQ

Qwen3 30B A3B — common questions

Is Qwen3 30B A3B on the AxForge serverless API?

Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.

What kind of model is Qwen3 30B A3B?

A mixture-of-experts LLM from the Qwen3 family with a 65,536-token context window.

How fast is it on your hardware?

AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card — the newer Qwen3.6 35B A3B has a community figure on its page.

What does a dedicated deployment look like?

A dedicated NVIDIA DGX Spark (GB10, 128 GB unified memory) rented by the hour, week, month or year, running Qwen3 30B A3B behind an OpenAI-compatible endpoint on your own machine, hosted in the EU with zero prompt retention.

What does it cost?

Two parts: the DGX Spark hardware rental — by the hour, week, month or year, with longer terms earning the lower rate — and the managed service, quoted per deployment. Both are confirmed in writing before anything is billed.

Ready for Qwen3 30B A3B on your own machine?

Request deployment Sign in Talk to an engineer

Explore

More from AxForge

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms