{"data":[{"id":"google/gemma-3-27b-it","name":"Google: Gemma 3 27B","created":1759532711,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":53920,"max_output_length":16384,"pricing":{"prompt":"0.00000015","completion":"0.00000046","input_cache_read":"0.000000075"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"hugging_face_id":"google/gemma-3-27b-it","is_tee":true,"providers":["near-ai"],"description":"Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to [Gemma 2](google/gemma-2-27b-it)","metadata":{}},{"id":"openai/gpt-oss-20b","name":"OpenAI: GPT OSS 20B","created":1759532709,"input_modalities":["text"],"output_modalities":["text"],"context_length":131072,"max_output_length":131072,"pricing":{"prompt":"0.00000004","completion":"0.00000015"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"openai/gpt-oss-20b","is_tee":true,"providers":["phala"],"description":"gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.","metadata":{"appid":"bfe88926c2826cf14a819ef0ae7558cac3bf024c"}},{"id":"qwen/qwen3-30b-a3b-instruct-2507","name":"Qwen3 30B A3B Instruct 2507","created":1764333076,"input_modalities":["text"],"output_modalities":["text"],"context_length":262144,"max_output_length":32000,"pricing":{"prompt":"0.00000015","completion":"0.00000055"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3-30B-A3B-Instruct-2507","deprecation_date":"2026-07-29T18:00:00Z","is_tee":true,"providers":["near-ai"],"description":"Qwen3-30B-A3B-Instruct-2507 is a mixture-of-experts (MoE) causal language model featuring 30.5 billion total parameters and 3.3 billion activated parameters per inference. It supports ultra-long context up to 262 K tokens and operates exclusively in non-thinking mode, delivering strong enhancements in instruction following, reasoning, logical comprehension, mathematics, coding, multilingual understanding, and alignment with user preferences.","metadata":{}},{"id":"deepseek/deepseek-chat-v3.1","name":"DeepSeek V3.1","created":1764333077,"input_modalities":["text"],"output_modalities":["text"],"context_length":163840,"max_output_length":32768,"pricing":{"prompt":"0.00000105","completion":"0.0000031"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"deepseek-ai/DeepSeek-V3.1","is_tee":true,"providers":["near-ai"],"description":"DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference.","metadata":{}},{"id":"meta-llama/llama-3.3-70b-instruct","name":"Meta: Llama 3.3 70B Instruct","created":1764320401,"input_modalities":["text"],"output_modalities":["text"],"context_length":131072,"max_output_length":16384,"pricing":{"prompt":"0.000002","completion":"0.000002"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"hugging_face_id":"meta-llama/Llama-3.3-70B-Instruct","is_tee":true,"providers":["tinfoil"],"description":"The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.","metadata":{}},{"id":"z-ai/glm-4.7","name":"Z.AI: GLM 4.7","created":1769662351,"input_modalities":["text"],"output_modalities":["text"],"context_length":131072,"max_output_length":131072,"pricing":{"prompt":"0.00000085","completion":"0.0000033"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"zai-org/GLM-4.7","deprecation_date":"2026-07-29T18:00:00Z","is_tee":true,"providers":["near-ai"],"description":"GLM-4.7 is Z.AI's latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while delivering more natural conversational experiences and superior front-end aesthetics.","metadata":{}},{"id":"phala/uncensored-24b","name":"Phala: Venice Uncensored 24B","created":1767996709,"input_modalities":["text"],"output_modalities":["text"],"context_length":32768,"max_output_length":32768,"pricing":{"prompt":"0.0000002","completion":"0.0000009"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"deprecation_date":"2026-07-29T18:00:00Z","is_tee":true,"providers":["phala"],"description":"Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving user control over alignment, system prompts, and behavior. Intended for advanced and unrestricted use cases, Venice Uncensored emphasizes steerability and transparent behavior, removing default safety and alignment layers typically found in mainstream assistant models.","metadata":{}},{"id":"minimax/minimax-m2.5","name":"MiniMax: MiniMax M2.5","created":1771685552,"input_modalities":["text"],"output_modalities":["text"],"context_length":196608,"max_output_length":196608,"pricing":{"prompt":"0.0000002","completion":"0.00000138","input_cache_read":"0.000000075"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"MiniMaxAI/MiniMax-M2.5","is_tee":true,"providers":["chutes"],"description":"MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams. Scoring 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp, M2.5 is also more token efficient than previous generations, having been trained to optimize its actions and output through planning.","metadata":{}},{"id":"qwen/qwen3.5-27b","name":"Qwen: Qwen3.5-27B","created":1773395704,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":65536,"pricing":{"prompt":"0.0000003","completion":"0.0000024","input_cache_read":"0.00000015"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3.5-27B","is_tee":true,"providers":["chutes"],"description":"The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.","metadata":{}},{"id":"z-ai/glm-5.2","name":"Z.ai: GLM 5.2","created":1781637446,"input_modalities":["text"],"output_modalities":["text"],"context_length":1048576,"max_output_length":131072,"pricing":{"prompt":"0.0000014","completion":"0.0000044","input_cache_read":"0.0000007"},"supported_parameters":["frequency_penalty","include_reasoning","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"hugging_face_id":"zai-org/GLM-5.2","is_tee":true,"providers":["chutes","near-ai","phala","tinfoil"],"description":"GLM-5.2 is Z.ai's flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context and execute long-running tasks more reliably. Served as a text-only TEE deployment via Phala.","metadata":{}},{"id":"deepseek/deepseek-v4-flash","name":"DeepSeek: DeepSeek V4 Flash","created":1780384051,"input_modalities":["text"],"output_modalities":["text"],"context_length":1048576,"max_output_length":65536,"pricing":{"prompt":"0.0000002","completion":"0.0000004"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"deepseek-ai/DeepSeek-V4-Flash","is_tee":true,"providers":["near-ai"],"description":"DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...","metadata":{}},{"id":"qwen/qwen3.6-27b","name":"Qwen: Qwen3.6 27B","created":1780549917,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":262144,"max_output_length":262140,"pricing":{"prompt":"0.00000032","completion":"0.0000027","input_cache_read":"0.00000015"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"Qwen/Qwen3.6-27B","is_tee":true,"providers":["chutes","near-ai"],"description":"Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities accepting text and image inputs, a configurable thinking/reasoning mode, and a native 262K context window. Served as a TEE deployment via Chutes.","metadata":{}},{"id":"openai/gpt-oss-120b","name":"OpenAI: GPT OSS 120B","created":1759532710,"input_modalities":["text"],"output_modalities":["text"],"context_length":131072,"max_output_length":131072,"pricing":{"prompt":"0.00000015","completion":"0.00000060"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"openai/gpt-oss-120b","is_tee":true,"providers":["near-ai","tinfoil"],"description":"gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.","metadata":{}},{"id":"phala/qwen3.6-35b-a3b-uncensored","name":"Phala: Qwen3.6 35B-A3B Uncensored (Aggressive)","created":1779562759,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":131072,"max_output_length":131072,"pricing":{"prompt":"0.0000003","completion":"0.0000015"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"is_tee":true,"providers":["phala"],"description":"Uncensored \"Aggressive\" variant of Qwen3.6-35B-A3B from Alibaba's Qwen team. The fine-tune by HauhauCS removes refusal behaviors (0/465 refusals) without modifying datasets or core capabilities. The base architecture is a 35B-parameter Mixture-of-Experts model with 256 experts routing 8 per token (~3B active params), 40 layers, and a hybrid linear+full-softmax attention mechanism (3:1 ratio). Supports a native 262K context and is natively multimodal across text, images, and video. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; FP8 quantization by lamianlbe.","metadata":{}},{"id":"moonshotai/kimi-k2.6","name":"MoonshotAI: Kimi K2.6","created":1776747825,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000109","completion":"0.0000046","input_cache_read":"0.00000037"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"moonshotai/Kimi-K2.6","is_tee":true,"providers":["chutes","tinfoil"],"description":"Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and demonstrates strong performance in agentic workflows.","metadata":{}},{"id":"qwen/qwen3.5-122b-a10b","name":"Qwen: Qwen3.5-122B-A10B","created":1779787333,"input_modalities":["text"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000046","completion":"0.00000368"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"Qwen/Qwen3.5-122B-A10B","is_tee":true,"providers":["near-ai"],"description":"Qwen3.5-122B-A10B is a large Mixture-of-Experts model from Alibaba Cloud with 122B total parameters and 10B active parameters per token. Strong on reasoning, coding, and tool calling with 262K context. Served as a text-only TEE deployment via NEAR AI.","metadata":{}},{"id":"qwen/qwen3-32b","name":"Qwen: Qwen3 32B","created":1779762873,"input_modalities":["text"],"output_modalities":["text"],"context_length":40960,"max_output_length":16384,"pricing":{"prompt":"0.00000012","completion":"0.0000005","input_cache_read":"0.000000052"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"Qwen/Qwen3-32B","is_tee":true,"providers":["chutes"],"description":"Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a thinking mode for complex tasks and a standard mode for general dialogue. Served as a TEE deployment via Chutes.","metadata":{}},{"id":"deepseek/deepseek-v4-pro","name":"DeepSeek: DeepSeek V4 Pro","created":1779762825,"input_modalities":["text"],"output_modalities":["text"],"context_length":800000,"max_output_length":384000,"pricing":{"prompt":"0.0000015","completion":"0.00000525"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"deepseek-ai/DeepSeek-V4-Pro","is_tee":true,"providers":["tinfoil"],"description":"DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and tool use.","metadata":{}},{"id":"qwen/qwen3.6-35b-a3b","name":"Qwen: Qwen3.6 35B A3B","created":1779762842,"input_modalities":["text"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.0000002","completion":"0.00000127"},"supported_parameters":[],"supported_sampling_parameters":[],"supported_features":[],"hugging_face_id":"Qwen/Qwen3.6-35B-A3B","is_tee":true,"providers":["near-ai"],"description":"Qwen3.6-35B-A3B is an open-weight model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated attention. Served as a text-only TEE deployment via NEAR AI.","metadata":{}},{"id":"google/gemma-4-31b-it","name":"Google: Gemma 4 31B","created":1779762858,"input_modalities":["text"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000015","completion":"0.00000046","input_cache_read":"0.000000075"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_a","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_a","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"hugging_face_id":"google/gemma-4-31B-it","is_tee":true,"providers":["chutes","near-ai","phala"],"description":"Gemma 4 31B Instruct is Google DeepMind's 30.7B dense model. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and strong multilingual performance. Served as a text-only TEE deployment via NEAR AI.","metadata":{}},{"id":"phala/gemma-4-26b-a4b-uncensored","name":"Phala: Gemma-4 26B-A4B Uncensored (Heretic)","created":1779562961,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":65536,"max_output_length":65536,"pricing":{"prompt":"0.00000015","completion":"0.0000007"},"supported_parameters":["frequency_penalty","logit_bias","logprobs","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","structured_outputs","tools"],"is_tee":true,"providers":["phala"],"description":"Uncensored \"Heretic\" variant of google/gemma-4-26B-A4B-it created using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method and row-norm preservation. Refusals drop from 100/100 to 11/100 with KL divergence 0.0499 vs the base model. The base Gemma 4 26B A4B is a Mixture-of-Experts model with 25.2B total / 3.8B active parameters (8 active / 128 total experts), 30-layer transformer with hybrid local sliding (1024) + global attention, supporting a 256K context window. Natively multimodal (text + images, variable aspect ratios). Strong on coding, reasoning, function calling, with native system prompt support across 35+ languages. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; vLLM-compatible FP8-Static quantization by cloud19 (router excluded from quantization).","metadata":{}},{"id":"z-ai/glm-5.1","name":"Z.ai: GLM 5.1","created":1776694000,"input_modalities":["text"],"output_modalities":["text"],"context_length":202752,"max_output_length":128000,"pricing":{"prompt":"0.00000121","completion":"0.0000042","input_cache_read":"0.0000006"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","parallel_tool_calls","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"zai-org/GLM-5.1","is_tee":true,"providers":["chutes","near-ai"],"description":"GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...","metadata":{}},{"id":"qwen/qwen3.5-397b-a17b","name":"Qwen: Qwen3.5 397B A17B","created":1772249193,"input_modalities":["text","image","video"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.00000055","completion":"0.0000035","input_cache_read":"0.000000225"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3.5-397B-A17B","is_tee":true,"providers":["chutes"],"description":"The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers state-of-the-art performance comparable to leading-edge models across a wide range of tasks, including language understanding, logical reasoning, code generation, agent-based tasks, image understanding, video understanding, and graphical user interface (GUI) interactions. With its robust code-generation and agent capabilities, the model exhibits strong generalization across diverse agent.","metadata":{}},{"id":"z-ai/glm-5","name":"Z.AI: GLM 5","created":1770707173,"input_modalities":["text"],"output_modalities":["text"],"context_length":202752,"max_output_length":202752,"pricing":{"prompt":"0.0000012","completion":"0.0000035","input_cache_read":"0.000000475"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"zai-org/GLM-5","is_tee":true,"providers":["chutes"],"description":"GLM-5 is an open-source foundation model built for complex systems engineering and long-horizon agent workflows. It delivers production-grade productivity for large-scale programming tasks, with performance aligned to top closed-source models, and is designed for expert developers building at the system level.","metadata":{}},{"id":"moonshotai/kimi-k2.5","name":"MoonshotAI: Kimi K2.5","created":1769672802,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":262144,"max_output_length":262144,"pricing":{"prompt":"0.0000006","completion":"0.000003","input_cache_read":"0.00000022"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"hugging_face_id":"moonshotai/Kimi-K2.5","is_tee":true,"providers":["chutes"],"description":"Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm. Built on Kimi K2 with continued pretraining over approximately 15T mixed visual and text tokens, it delivers strong performance in general reasoning, visual coding, and agentic tool-calling.","metadata":{}},{"id":"deepseek/deepseek-v3.2","name":"DeepSeek: DeepSeek V3.2","created":1764804811,"input_modalities":["text"],"output_modalities":["text"],"context_length":163840,"max_output_length":64000,"pricing":{"prompt":"0.000001","completion":"0.000001","input_cache_read":"0.0000005"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","max_tokens","min_p","presence_penalty","reasoning","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","reasoning","structured_outputs","tools"],"hugging_face_id":"deepseek-ai/DeepSeek-V3.2","is_tee":true,"providers":["chutes"],"description":"DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.","metadata":{}},{"id":"qwen/qwen3-vl-30b-a3b-instruct","name":"Qwen: Qwen3 VL 30B A3B Instruct","created":1764320402,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":128000,"max_output_length":32768,"pricing":{"prompt":"0.0000002","completion":"0.0000007"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","structured_outputs","tools"],"hugging_face_id":"Qwen/Qwen3-VL-30B-A3B-Instruct","is_tee":true,"providers":["near-ai"],"description":"Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception of real-world/synthetic categories, 2D/3D spatial grounding, and long-form visual comprehension, achieving competitive multimodal benchmark results. For agentic use, it handles multi-image multi-turn instructions, video timeline alignments, GUI automation, and visual coding from sketches to debugged UI. Text performance matches flagship Qwen3 models, suiting document AI, OCR, UI assistance, spatial tasks, and agent research.","metadata":{}},{"id":"qwen/qwen-2.5-7b-instruct","name":"Qwen2.5 7B Instruct","created":1759532719,"input_modalities":["text"],"output_modalities":["text"],"context_length":32768,"max_output_length":32768,"pricing":{"prompt":"0.00000004","completion":"0.0000001"},"supported_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","response_format","seed","stop","temperature","tool_choice","tools","top_k","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","tools"],"hugging_face_id":"Qwen/Qwen2.5-7B-Instruct","is_tee":true,"providers":["phala"],"description":"Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2:\n\n- Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains.\n\n- Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and condition-setting for chatbots.\n\n- Long-context Support up to 128K tokens and can generate up to 8K tokens.\n\n- Multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more.\n\nUsage of this model is subject to [Tongyi Qianwen LICENSE AGREEMENT](https://huggingface.co/Qwen/Qwen1.5-110B-Chat/blob/main/LICENSE).","metadata":{"appid":"0f765d7d4b1005bbfb7e77f710d6776bfb11f3ab"}},{"id":"moonshotai/kimi-k3","name":"MoonshotAI: Kimi K3","created":1785294047,"input_modalities":["text","image"],"output_modalities":["text"],"context_length":1048576,"max_output_length":65535,"pricing":{"prompt":"0.000003","completion":"0.000015","input_cache_read":"0.0000015"},"supported_parameters":["frequency_penalty","include_reasoning","logit_bias","logprobs","max_tokens","min_p","presence_penalty","reasoning","reasoning_effort","repetition_penalty","response_format","seed","stop","structured_outputs","temperature","tool_choice","tools","top_k","top_logprobs","top_p"],"supported_sampling_parameters":["frequency_penalty","logit_bias","max_tokens","min_p","presence_penalty","repetition_penalty","seed","stop","temperature","top_k","top_p"],"supported_features":["json_mode","logprobs","reasoning","structured_outputs","tools"],"is_tee":true,"providers":["chutes"],"description":"Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows.","metadata":{}}]}