Skip to content
Qwen logo

qwen3.8-27b

Image-Text-to-Text • Qwen

View as MarkdownAgent setup
  • Cloudflare-hosted
  • Function calling
  • Reasoning
  • Vision

Qwen 3.8 27B is a 27-billion-parameter instruction-tuned language model from Alibaba's Qwen family, designed for vision, efficient general-purpose text generation and agentic workloads.

Model Info
Context Window262,144 tokens
Function calling Yes
ReasoningYes
VisionYes
Unit Pricing$0.45 per M input tokens, $3.20 per M output tokens

Parameters

Synchronous — Send a request and receive a complete response
Input format
prompt
stringrequiredminLength: 1The input text prompt for the model to generate a response.
model
stringID of the model to use (e.g. '@cf/zai-org/glm-4.7-flash, etc').
frequency_penalty
number | nullPenalizes new tokens based on their existing frequency in the text so far.
logit_bias
object | nullModify the likelihood of specified tokens appearing in the completion. Maps token IDs to bias values from -100 to 100.
logprobs
boolean | nullWhether to return log probabilities of the output tokens.
top_logprobs
integer | nullHow many top log probabilities to return at each token position (0-20). Requires logprobs=true.
max_tokens
integer | nullDeprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
max_completion_tokens
integer | nullAn upper bound for the number of tokens that can be generated for a completion.
metadata
object | nullSet of 16 key-value pairs that can be attached to the object.
modalities
array | nullOutput types requested from the model (e.g. ['text'] or ['text', 'audio']).
n
integer | nullHow many chat completion choices to generate for each input message.
parallel_tool_calls
booleandefault: trueWhether to enable parallel function calling during tool use.
presence_penalty
number | nullPenalizes new tokens based on whether they appear in the text so far.
reasoning_effort
string | nullConstrains effort on reasoning for reasoning models (o1, o3-mini, etc.).
seed
integer | nullIf specified, the system will make a best effort to sample deterministically.
service_tier
string | nullSpecifies the processing type used for serving the request.
store
boolean | nullWhether to store the output for model distillation / evals.
stream
boolean | nullIf true, partial message deltas will be sent as server-sent events.
temperature
number | nullSampling temperature between 0 and 2.
top_p
number | nullNucleus sampling: considers the results of the tokens with top_p probability mass.
user
stringA unique identifier representing your end-user, for abuse monitoring.
id
stringA unique identifier for the chat completion.
object
string
created
integerUnix timestamp (seconds) of when the completion was created.
model
stringThe model used for the chat completion.
system_fingerprint
string | null
service_tier
string | null
Streaming — Send a request with `stream: true` and receive server-sent events
Input format
prompt
stringrequiredminLength: 1The input text prompt for the model to generate a response.
model
stringID of the model to use (e.g. '@cf/zai-org/glm-4.7-flash, etc').
frequency_penalty
number | nullPenalizes new tokens based on their existing frequency in the text so far.
logit_bias
object | nullModify the likelihood of specified tokens appearing in the completion. Maps token IDs to bias values from -100 to 100.
logprobs
boolean | nullWhether to return log probabilities of the output tokens.
top_logprobs
integer | nullHow many top log probabilities to return at each token position (0-20). Requires logprobs=true.
max_tokens
integer | nullDeprecated in favor of max_completion_tokens. The maximum number of tokens to generate.
max_completion_tokens
integer | nullAn upper bound for the number of tokens that can be generated for a completion.
metadata
object | nullSet of 16 key-value pairs that can be attached to the object.
modalities
array | nullOutput types requested from the model (e.g. ['text'] or ['text', 'audio']).
n
integer | nullHow many chat completion choices to generate for each input message.
parallel_tool_calls
booleandefault: trueWhether to enable parallel function calling during tool use.
presence_penalty
number | nullPenalizes new tokens based on whether they appear in the text so far.
reasoning_effort
string | nullConstrains effort on reasoning for reasoning models (o1, o3-mini, etc.).
seed
integer | nullIf specified, the system will make a best effort to sample deterministically.
service_tier
string | nullSpecifies the processing type used for serving the request.
store
boolean | nullWhether to store the output for model distillation / evals.
stream
boolean | nullIf true, partial message deltas will be sent as server-sent events.
temperature
number | nullSampling temperature between 0 and 2.
top_p
number | nullNucleus sampling: considers the results of the tokens with top_p probability mass.
user
stringA unique identifier representing your end-user, for abuse monitoring.
type
string
contentType
text/event-stream
format
binary
Batch — Send multiple requests in a single API call
id
stringA unique identifier for the chat completion.
object
string
created
integerUnix timestamp (seconds) of when the completion was created.
model
stringThe model used for the chat completion.
system_fingerprint
string | null
service_tier
string | null

API Schemas (Raw)

SynchronousInput
SynchronousOutput
StreamingInput
StreamingOutput
BatchInput
BatchOutput

Was this helpful?