Dedicated model · Available as managed deployment
The sentence-transformers workhorse — a 23M-parameter embedding model producing 384-dimensional vectors, the most downloaded embedding model in the world, under Apache-2.0. Validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only — an OpenAI-compatible endpoint on hardware only you use, operated by AxForge in the EU.
Why AxForge
| The most used embedding model | Hundreds of millions of downloads: the default in tutorials, vector databases and RAG starters everywhere. |
|---|---|
| Tiny and fast | 23M parameters and 384 dimensions — thousands of embeddings per second on a single machine, cheap to store and search. |
| Apache-2.0 | Permissive — commercial use, no licence conversation. |
Specifications
| Model | all-MiniLM-L6-v2 — sentence-transformers |
|---|---|
| Modalities | Text → vector (embeddings) |
| Sizes | 23M |
| Context window | 512 tokens |
| Licence | Open weights — apache-2.0; commercial use permitted |
| Hardware | NVIDIA DGX Spark (GB10, 128 GB unified memory) — owned and operated by AxForge |
| Rental term | Hour, week, month or year |
| Hardware pricing | €0.69/hour on demand · €0.66/hour by the week · €0.62/hour by the month · €0.55/hour by the year, excl. VAT |
| Managed service | Quoted per deployment |
| Region | Málaga, Spain (eu-es-1) |
Full details, benchmarks and FAQ on the all-MiniLM-L6-v2 page. Prices exclude VAT.
How it works
| 1 | Request deployment — describe your traffic, context needs and rental term. |
|---|---|
| 2 | You receive the configuration, hardware rental and managed-service price in writing before anything is billed. |
| 3 | AxForge deploys all-MiniLM-L6-v2 on a dedicated DGX Spark reserved for you. |
| 4 | Point your OpenAI SDK at your own endpoint with the model name you receive. |
| 5 | Adjust the term — hour, week, month or year — as your workload settles. |
Request deployment or sign in to start.
FAQ
Not on the serverless API — it is available as a managed deployment: validated on AxForge hardware and deployed on a dedicated DGX Spark for your traffic only. The serverless API serves Qwen3.8 27B.
English text, short passages and high volume where speed and cost matter more than the last few points of retrieval quality; for multilingual or long inputs look at BGE-M3 or Qwen3 Embedding.
Yes — a dedicated DGX Spark runs a small embedding model next to an LLM behind one OpenAI-compatible endpoint.
AxForge publishes only numbers it measures itself, and has not benchmarked this model on its nodes yet. For quality benchmarks, see the official model card.
Hardware by the hour, week, month or year; the managed service is quoted per deployment — both confirmed in writing before anything is billed.