For most new applications, prefer the Chat Completions endpoint. It supports more capable models, structured conversations, and function calling. The completions endpoint exists primarily for backward compatibility.
Endpoint
Request parameters
string
required
The model to use. Use the provider-prefixed format (
openai/gpt-3.5-turbo-instruct) or the short name where unambiguous. Retrieve available model IDs from GET /v1/models.string | string[]
required
The prompt text to complete. Pass a string for a single prompt or an array of strings to generate completions for multiple prompts in one request.
integer
default:"16"
Maximum number of tokens to generate per completion.
number
default:"1"
Sampling temperature between
0 and 2. Lower values are more deterministic; higher values are more creative.boolean
default:"false"
Stream the response as server-sent events. Each event contains a partial completion delta. The stream ends with
data: [DONE].string | string[]
One or more stop sequences. Generation stops when any sequence is encountered; the stop sequence itself is not included in the output.
integer
default:"1"
Generate this many completions server-side and return the best one (as measured by log probability). Higher values increase latency and token usage.
integer
Include log probabilities for the top
logprobs tokens at each position. Maximum value is 5.Response fields
string
Unique identifier for this completion request.
string
Always
"text_completion".string
The model that served the request.
object[]
object