No models found
Try a different search term, or broaden your search by removing filters.
bge-base-en-v1.5
BAAI general embedding (Base) model that transforms any given text into a 768-dimensional vector
- Cloudflare-hosted
- Batch
- Context: 153.6K tokens
- Pricing listed
bge-m3
Multi-Functionality, Multi-Linguality, and Multi-Granularity embeddings model.
- Cloudflare-hosted
- Context: 60K tokens
- Pricing listed
deepseek-r1-distill-qwen-32b
DeepSeek-R1-Distill-Qwen-32B is a model distilled from DeepSeek-R1 based on Qwen2.5. It outperforms OpenAI-o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.
- Cloudflare-hosted
- Reasoning
- Context: 80K tokens
- Pricing listed
deepseek-v4-flash-0731
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Context: 1.3M tokens
- Pricing listed
deepseek-v4-pro-0813
DeepSeek V4 Pro is a high-capability reasoning model from DeepSeek with a one million token context window, built for long-horizon agentic workflows and complex, multi-step problem-solving
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 1M tokens
- Pricing listed
gemma-4-26b-a4b-it
Gemma 4 is Google's most intelligent family of open models, built from Gemini 3 research to maximize intelligence-per-parameter.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Vision
- Context: 256K tokens
- Pricing listed
gemma-sea-lion-v4-27b-it
SEA-LION stands for Southeast Asian Languages In One Network, which is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
- Cloudflare-hosted
- Context: 128K tokens
- Pricing listed
glm-4.7-flash
GLM-4.7-Flash is a fast and efficient multilingual text generation model with a 131,072 token context window. Optimized for dialogue, instruction-following, and multi-turn tool calling across 100+ languages.
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 131.1K tokens
- Pricing listed
glm-5.2
Z.ai's flagship agentic coding model
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 262.1K tokens
- Pricing listed
glm-5.3
GLM-5.3 is Z.ai's flagship agentic coding model, pairing a 1M-token context window with reasoning, function calling, and structured outputs to power multi-step, tool-driven development workflows.
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 1.3M tokens
- Pricing listed
gpt-oss-120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases – gpt-oss-120b is for production, general purpose, high reasoning use-cases.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Context: 128K tokens
- Pricing listed
gpt-oss-20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases – gpt-oss-20b is for lower latency, and local or specialized use-cases.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Context: 128K tokens
- Pricing listed
kimi-k2.6
Kimi K2.6 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Vision
- Context: 262.1K tokens
- Pricing listed
kimi-k2.7-code
Kimi K2.7 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
- Cloudflare-hosted
- Function calling
- Reasoning
- Vision
- Context: 262.1K tokens
- Pricing listed
llama-3.1-8b-instruct-fp8
Llama 3.1 8B quantized to FP8 precision
- Cloudflare-hosted
- Context: 32K tokens
- Pricing listed
llama-3.2-11b-vision-instruct
The Llama 3.2-Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image.
- Cloudflare-hosted
- LoRA
- Vision
- Context: 128K tokens
- Pricing listed
llama-3.2-1b-instruct
The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks.
- Cloudflare-hosted
- Context: 60K tokens
- Pricing listed
llama-3.2-3b-instruct
The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks.
- Cloudflare-hosted
- LoRA
- Context: 80K tokens
- Pricing listed
llama-3.3-70b-instruct-fp8-fast
Llama 3.3 70B quantized to fp8 precision, optimized to be faster.
- Cloudflare-hosted
- Batch
- Function calling
- Context: 24K tokens
- Pricing listed
melotts
MeloTTS is a high-quality multi-lingual text-to-speech library by MyShell.ai.
- Cloudflare-hosted
- Pricing listed
nemotron-3-120b-a12b
NVIDIA Nemotron 3 Super is a hybrid MoE model with leading accuracy for multi-agent applications and specialized agentic AI systems.
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 256K tokens
- Pricing listed
qwen3-embedding-0.6b
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks.
- Cloudflare-hosted
- Context: 8.2K tokens
- Pricing listed
qwen3.8-27b
Qwen 3.8 27B is a 27-billion-parameter instruction-tuned language model from Alibaba's Qwen family, designed for vision, efficient general-purpose text generation and agentic workloads.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Vision
- Context: 262.1K tokens
- Pricing listed