Endpoint
Request parameters
string
required
The embedding model to use. For example:
text-embedding-3-small or text-embedding-3-large. Retrieve the full list of available embedding models from GET /v1/models.string | string[]
required
The text to embed. Pass a single string for one embedding or an array of strings to embed multiple inputs in a single request. Arrays are more efficient than making one request per string.
string
default:"float"
Format for the returned vectors. Use
"float" for a JSON array of numbers (default), or "base64" for a base64-encoded binary representation that reduces response size.integer
Number of dimensions in the output embedding. Only supported by models that accept a dimensions parameter (e.g.,
text-embedding-3-small, text-embedding-3-large). Reducing dimensions trades some accuracy for lower storage and compute cost.Response fields
string
Always
"list".object[]
string
The model that generated the embeddings.
object
Examples
Single string input
Array input (batch embedding)
Common use cases
- RAG pipelines — embed your documents at index time, embed user queries at runtime, and retrieve the closest chunks by cosine similarity before passing them to a chat completions request.
- Semantic search — find documents that are conceptually similar to a query even when no keywords match.
- Document clustering — group large collections of text by topic without labeled training data.
- Classification — train a lightweight classifier on top of embeddings rather than fine-tuning a full model.