> For the complete documentation index, see [llms.txt](https://docs.nebulablock.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.nebulablock.com/products/serverless-inference/model-catalog.md).

# Model Catalog

Every model available on Nebula Block serverless inference, with model IDs and context lengths, grouped by capability.

Nebula Block serves every model below through one OpenAI-compatible endpoint at `https://inference.nebulablock.com/v1`, grouped here by what the model does.

Use the **Model ID** column as the `model` parameter in your API calls.

> **Note:** The catalog changes frequently. For a machine-readable list that is always current, call [`GET /v1/models`](/api-reference/inference-api/list-models.md), or browse [Serverless Models](https://console.nebulablock.com/serverless) in the console. Per-token pricing — including promotional rates — is published on the [pricing page](https://www.nebulablock.com/pricing/serverless-ai) and in the console, so this page deliberately does not duplicate it.

Models marked 🇨🇦 are hosted in Canada — relevant if you have data-residency requirements. See [Sovereign AI Compute](https://www.nebulablock.com/ai-sovereign-compute).

## Text generation and multimodal chat

Called through [Chat Completions](/api-reference/inference-api/chat-completions.md). Models listed as multimodal accept images alongside text in the same request.

| Model                                        | Model ID                                        | Context | Description                                                                                                    |
| -------------------------------------------- | ----------------------------------------------- | ------- | -------------------------------------------------------------------------------------------------------------- |
| **GLM-5.3**                                  | `zai-org/GLM-5.3`                               | 1M      | Z.ai's GLM-5.3, a 1M-token-context reasoning model with tool calling and structured outputs. Reasoning is…     |
| **GLM-5.3-Flash**                            | `zai-org/GLM-5.3-Flash`                         | 1M      | Z.ai's fast, low-cost GLM-5.3 variant with a 1M-token context window, image understanding and tool calling.…   |
| **Qwen3.8-27B**                              | `Qwen/Qwen3.8-27B`                              | 64K     | Qwen's compact 27B dense model from the Qwen3.8 generation, with strong agentic and coding performance, tool…  |
| **Gemini-3.7-Flash**                         | `gemini/gemini-3.7-flash`                       | 1M      | Google's newest Flash model — stronger multimodal reasoning, coding and agentic performance at Flash speed,…   |
| **Qwen3.8-2.4T-A95B**                        | `Qwen/Qwen3.8-2.4T-A95B`                        | 256K    | Alibaba's most capable open-weight model — a 2.4T-parameter Mixture-of-Experts with 95B active parameters per… |
| **DeepSeek-V4-Pro-0813**                     | `deepseek-ai/DeepSeek-V4-Pro-0813`              | 1M      | DeepSeek's GA release of V4 Pro, a large-scale mixture-of-experts model with a 1M-token context window and…    |
| **Grok-4.6**                                 | `x-ai/grok-4.6`                                 | 500K    | xAI's smartest model, with frontier performance on coding, knowledge work, and STEM. 500K-token context…       |
| **DeepSeek-V4-Flash-0731**                   | `deepseek-ai/DeepSeek-V4-Flash-0731`            | 1M      | DeepSeek's official V4 Flash release, superseding the preview with substantially stronger agentic and coding…  |
| **Claude-Opus-5**                            | `anthropic/claude-opus-5`                       | 1M      | Anthropic's most capable model for complex agentic coding and enterprise work, delivering a step-change in…    |
| **Gemini-3.6-Flash**                         | `gemini/gemini-3.6-flash`                       | 1M      | Google's latest Flash model — upgraded multimodal reasoning, coding and agentic performance at Flash speed…    |
| **Gemini-3.5-Flash-Lite**                    | `gemini/gemini-3.5-flash-lite`                  | 1M      | Google's most cost-efficient Gemini 3.5 model — low-latency multimodal inference tuned for high-volume…        |
| **Qwen3.8-Max-Preview**                      | `Qwen/Qwen3.8-Max-Preview`                      | 991K    | Alibaba Qwen 2.4T-parameter flagship preview with major gains in coding, full-stack development, data…         |
| **Kimi-K3**                                  | `moonshotai/Kimi-K3`                            | 1M      | Moonshot AI's flagship model with a 1M-token context window, built for long-horizon agentic coding and tool…   |
| **GPT-5.6-Sol**                              | `openai/gpt-5.6-sol`                            | 1.05M   | OpenAI GPT-5.6 Sol — flagship multimodal model in the GPT-5.6 series with top-tier reasoning and agentic…      |
| **GPT-5.6-Terra**                            | `openai/gpt-5.6-terra`                          | 1.05M   | OpenAI GPT-5.6 Terra — a balanced multimodal model in the GPT-5.6 series with strong reasoning, positioned…    |
| **Grok-4.5**                                 | `x-ai/grok-4.5`                                 | 500K    | xAI's flagship model released in July 2026, with frontier performance on coding, knowledge work, and STEM.…    |
| **Claude-Sonnet-5**                          | `anthropic/claude-sonnet-5`                     | 1M      | Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use,… |
| **Claude-Fable-5**                           | `anthropic/claude-fable-5`                      | 1M      | Anthropic's most capable frontier model for demanding reasoning and long-horizon agentic work, with always-on… |
| **MiniMax-M3**                               | `MiniMaxAI/MiniMax-M3`                          | 1M      | A million-token multimodal frontier model from MiniMax, built for long-horizon agentic work with native…       |
| **Qwen3.6-35B-A3B** 🇨🇦                     | `Qwen/Qwen3.6-35B-A3B`                          | 256K    | Qwen3.6 35B (3B-active) MoE with built-in NEXTN speculative decoding, reasoning and tool calling — fast…       |
| **Claude-Opus-4.8**                          | `anthropic/claude-opus-4-8`                     | 1M      | Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and…      |
| **Gemini-3.5-Flash**                         | `gemini/gemini-3.5-flash`                       | 1M      | Latest Gemini 3.5 Flash — fast multimodal reasoning with 1M context                                            |
| **Gemini-3.1-Flash-Lite**                    | `gemini/gemini-3.1-flash-lite`                  | 1M      | Google's most cost-efficient, low-latency multimodal model in the Gemini 3 series, optimized for high-volume,… |
| **Grok-4.3**                                 | `x-ai/grok-4.3`                                 | 1M      | xAI's reasoning model released in late April 2026, featuring always-on reasoning, a 1-million-token context…   |
| **GPT-5.5**                                  | `openai/gpt-5.5`                                | 400K    | OpenAI's next-generation multimodal GPT-5.5 model with improved token efficiency for hard reasoning, coding,…  |
| **DeepSeek-V4-Pro**                          | `deepseek-ai/DeepSeek-V4-Pro`                   | 1M      | DeepSeek's next-generation flagship language model delivering state-of-the-art reasoning, coding, and agentic… |
| **DeepSeek-V4-Flash**                        | `deepseek-ai/DeepSeek-V4-Flash`                 | 1M      | DeepSeek's fast, cost-efficient V4 variant optimized for high-throughput reasoning and coding with a 1M…       |
| **MiMo-V2.5-Pro**                            | `XiaomiMiMo/MiMo-V2.5-Pro`                      | 1M      | Xiaomi's flagship text model tuned for complex software engineering, agentic workflows, and long-horizon…      |
| **MiMo-V2.5**                                | `XiaomiMiMo/MiMo-V2.5`                          | 1M      | Xiaomi's natively omnimodal model accepting text, image, audio, and video inputs with a 1M context window for… |
| **Gemma-4-31B**                              | `google/gemma-4-31b-it`                         | 128K    | Google DeepMind's open-weights, instruction-tuned multimodal model optimized for reasoning, coding, and…       |
| **Claude-Opus-4.7**                          | `anthropic/claude-opus-4-7`                     | 1M      | Anthropic's most capable generally available model, excelling at agentic coding, long-running tasks, and…      |
| **MiMo-V2-Omni**                             | `XiaomiMiMo/MiMo-V2-Omni`                       | 256K    | Xiaomi's omni-modal agent model that natively understands text, images, video, and audio in a unified…         |
| **Mistral-Small-4**                          | `mistralai/Mistral-Small-4-119B-2603`           | 256K    | A large (\~119B parameter) dense language model designed for efficient, high-quality text generation and…      |
| **Gemini-3.1-Flash-Lite-Preview**            | `gemini/gemini-3.1-flash-lite-preview`          | 1M      | Google’s fastest and most cost-efficient Gemini 3 model, optimized for high-throughput tasks and scalable…     |
| **Gemini-3.1-Pro-Preview**                   | `gemini/gemini-3.1-pro-preview`                 | 1M      | Google's frontier reasoning model that builds on the Gemini 3 Pro series with enhanced thinking capabilities,… |
| **MiMo-V2-Pro**                              | `XiaomiMiMo/MiMo-V2-Pro`                        | 1M      | A trillion-parameter text-only flagship agent model from Xiaomi, built for complex coding, long-horizon…       |
| **MiniMax-M2.7**                             | `MiniMaxAI/MiniMax-M2.7`                        | 200K    | A next-generation agentic LLM with native interleaved thinking, built for complex real-world productivity…     |
| **GLM-4.7-Flash**                            | `zai-org/GLM-4.7-Flash`                         | 128K    | A 30B Mixture-of-Experts language model optimized for efficient inference, strong coding performance, and…     |
| **Qwen3.5-Plus**                             | `Qwen/Qwen3.5-Plus`                             | 1M      | Alibaba's native vision-language model with a 1M-token context window, hybrid MoE architecture, and built-in…  |
| **Kimi-K2.5**                                | `moonshotai/Kimi-K2.5`                          | 256K    | An open-source native multimodal AI model that combines vision and language understanding with advanced…       |
| **Claude-Sonnet-4.6**                        | `anthropic/claude-sonnet-4-6`                   | 1M      | Anthropic's most capable mid-tier model, delivering near-Opus-level intelligence across coding, computer use,… |
| **Claude-Opus-4.6**                          | `anthropic/claude-opus-4-6`                     | 1M      | Anthropic's most intelligent model, excelling at complex reasoning, coding, analysis, and multi-step tasks     |
| **Claude-Opus-4.5**                          | `anthropic/claude-opus-4-5-20251101`            | 200K    | An advanced flagship model delivering top-tier reasoning, deep analysis, and long-context performance for the… |
| **Claude-Haiku-4.5**                         | `anthropic/claude-haiku-4-5-20251001`           | 200K    | A lightweight, low-latency model built for rapid responses, simple reasoning, and high-throughput applications |
| **Claude-Sonnet-4.5**                        | `anthropic/claude-sonnet-4-5-20250929`          | 1M      | A fast, cost-efficient model optimized for everyday reasoning, coding, and high-quality conversational tasks   |
| **Claude-Opus-4**                            | `anthropic/claude-opus-4-20250514`              | 200K    | Top-tier, large-scale reasoning model designed for complex analysis, long-context understanding, and…          |
| **Grok-4-Fast**                              | `x-ai/grok-4-fast`                              | 2M      | High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation |
| **Gemini-3-Flash-Preview**                   | `gemini/gemini-3-flash-preview`                 | 1M      | Google’s agentic workhorse model, bringing near Pro agentic, coding and multimodal intelligence, with more…    |
| **Gemini-3-Pro-Preview**                     | `gemini/gemini-3-pro-preview`                   | 1M      | Google’s latest multimodal AI model that combines advanced reasoning, coding, and image-video understanding…   |
| **Gemini-2.5-Pro**                           | `gemini/gemini-2.5-pro`                         | 1M      | Strongest Gemini, 1M context, great for code & knowledge                                                       |
| **Gemini-2.5-Flash**                         | `gemini/gemini-2.5-flash`                       | 1M      | Fast reasoning with improved latency & accuracy                                                                |
| **Gemini-2.5-Flash-Lite**                    | `gemini/gemini-2.5-flash-lite`                  | 1M      | Ultra-low latency Gemini variant                                                                               |
| **GPT-5.4**                                  | `openai/gpt-5.4`                                | 400K    | A frontier large language model optimized for complex professional tasks, combining strong reasoning, coding,… |
| **GPT-5.3-Chat**                             | `openai/gpt-5.3-chat`                           | 400K    | A fast, general-purpose conversational LLM based on the GPT-5.3 architecture, optimized for high-quality…      |
| **GPT-5.3-Codex**                            | `openai/gpt-5.3-codex`                          | 400K    | A high-performance agentic coding model designed to autonomously write, debug, and manage complex software…    |
| **GPT-5.2**                                  | `openai/gpt-5.2`                                | 400K    | OpenAI’s flagship multimodal (text + image) GPT-5.2 model for top-tier coding and agentic tasks, producing…    |
| **GPT-5.1**                                  | `openai/gpt-5.1`                                | 400K    | OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across…         |
| **GPT-5**                                    | `openai/gpt-5`                                  | 400K    | OpenAI’s multimodal (text + image) GPT-5 model for strong coding, reasoning, and agentic tasks across…         |
| **GPT-5-Mini**                               | `openai/gpt-5-mini`                             | 400K    | OpenAI’s faster, more cost-efficient GPT-5 model for well-defined tasks, supporting text and image input with… |
| **GPT-5-Nano**                               | `openai/gpt-5-nano`                             | 400K    | OpenAI’s fastest, most cost-efficient GPT-5 model for lightweight tasks like summarization and…                |
| **GPT-4o-mini**                              | `openai/gpt-4o-mini`                            | 125K    | Compact GPT-4 Omni, supports text & image inputs                                                               |
| **GLM-5**                                    | `zai-org/GLM-5`                                 | 198K    | Z.ai's most powerful model — a 744B-parameter (40B active) open-source MoE optimized for reasoning, coding,…   |
| **GLM-4.7**                                  | `zai-org/GLM-4.7`                               | 198K    | An open-weights MoE chat model from Z.ai optimized for strong reasoning and tool-using agents                  |
| **Kimi-K2-Thinking**                         | `moonshotai/Kimi-K2-Thinking`                   | 256K    | 1T-parameter reasoning-focused language model designed for complex, multi-step problem solving and tool use    |
| **DeepSeek-V3.2**                            | `deepseek-ai/DeepSeek-V3.2`                     | 160K    | A high-performance open-weight large language model from DeepSeek optimized for strong reasoning, coding, and… |
| **DeepSeek-V3.2-Exp**                        | `deepseek-ai/DeepSeek-V3.2-Exp`                 | 160K    | Experimental version with sparse-attention architecture for long-context efficiency                            |
| **DeepSeek-V3.1**                            | `deepseek-ai/DeepSeek-V3.1`                     | 160K    | Hybrid inference LLM with Think/Non-Think modes, 128K context, advanced agent                                  |
| **DeepSeek-V3-0324**                         | `deepseek-ai/DeepSeek-V3-0324`                  | 64K     | High-performance open-source MoE language model optimized for reasoning, coding, and efficient text generation |
| **DeepSeek-R1-0528**                         | `deepseek-ai/DeepSeek-R1-0528`                  | 160K    | An open-source next-generation reasoning-optimized language model with enhanced logic, math, and code…         |
| **Qwen3-235B-A22B-Instruct-2507**            | `Qwen/Qwen3-235B-A22B-Instruct-2507`            | 256K    | A 235B-parameter MoE instruction-tuned language model for general, multilingual, and coding tasks              |
| **Mistral-Small-3.2-24B-Instruct-2506** 🇨🇦 | `mistralai/Mistral-Small-3.2-24B-Instruct-2506` | 32K     | 24B instruction model, long context, fewer errors                                                              |
| **Llama3.3-70B**                             | `meta-llama/Llama-3.3-70B-Instruct`             | 38K     | Multilingual 70B delivering 405B-level performance                                                             |
| **Kimi-K2.6**                                | `moonshotai/Kimi-K2.6`                          | 256K    | An open-weight 1T-parameter MoE multimodal model built for long-horizon agentic coding, with a 256K context…   |

## Vision

Dedicated vision-language models for image understanding. See [Vision](/products/serverless-inference/vision.md).

| Model                      | Model ID                      | Context | Description                                                                                                   |
| -------------------------- | ----------------------------- | ------- | ------------------------------------------------------------------------------------------------------------- |
| **Qwen3-VL-Plus**          | `Qwen/Qwen3-VL-Plus`          | 256K    | Alibaba's multimodal vision-language API model that handles text, image, and video inputs with strong visual… |
| **Qwen3-VL-Flash**         | `Qwen/Qwen3-VL-Flash`         | 256K    | Alibaba's lightweight, cost-effective multimodal vision-language API model designed for fast inference on…    |
| **Qwen2.5-VL-7B-Instruct** | `Qwen/Qwen2.5-VL-7B-Instruct` | 125K    | Vision-language model for multimodal understanding                                                            |

## Image generation

Called through [Images](/api-reference/inference-api/images.md). See [Image Generation](/products/serverless-inference/image-generation.md).

| Model                      | Model ID                                | Modalities                         | Description                                                                                                 |
| -------------------------- | --------------------------------------- | ---------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| **Nano-Banana-2**          | `gemini/gemini-3.1-flash-image-preview` | text\_to\_image, image\_generation | Text-to-image and image+text-to-image with up to 14 reference images. Supports aspect\_ratio, image\_size,… |
| **Nano-Banana-Pro-Edit**   | `gemini/gemini-3-pro-image-edit`        | image\_to\_image, text\_to\_image  | Premium AI-powered image editing with Gemini 3 Pro. Advanced editing capabilities with better quality,…     |
| **Nano-Banana-Pro**        | `gemini/gemini-3-pro-image-preview`     | text\_to\_image                    | High-quality image generation with better text rendering, character consistency, and advanced composition.… |
| **Nano-Banana-Edit**       | `gemini/gemini-2.5-flash-image-edit`    | image\_to\_image, text\_to\_image  | AI-powered image editing with Gemini 2.5 Flash. Edit images using natural language instructions - remove…   |
| **Bytedance-Seedream-3.0** | `Bytedance/seedream-3-0-t2i-250415`     | text\_to\_image, image\_to\_image  | Bilingual text-to-image, 2K resolution, accurate text & artistic layouts                                    |

## Video generation

Called through [Videos](/api-reference/inference-api/videos.md). See [Video Generation](/products/serverless-inference/video-generation.md).

| Model                                | Model ID                                | Modalities                                          | Description                                                                                                  |
| ------------------------------------ | --------------------------------------- | --------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| **Seedance-2.0**                     | `Byteplus/seedance-2-0-260128`          | text\_to\_video, image\_to\_video, video\_to\_video | BytePlus Seedance 2.0 — unified multimodal video (text/image/video/audio refs, native audio). Billed on…     |
| **Seedance-2.0-Fast**                | `Byteplus/seedance-2-0-fast-260128`     | text\_to\_video, image\_to\_video, video\_to\_video | BytePlus Seedance 2.0 Fast — faster/cheaper variant (no 1080p). Billed on actual tokens.                     |
| **Seedance-2.0-Mini**                | `Byteplus/seedance-2-0-mini-260615`     | text\_to\_video, image\_to\_video, video\_to\_video | BytePlus Seedance 2.0 Mini — cost-effective variant (no 1080p). Billed on actual tokens.                     |
| **Veo-3.1-Fast**                     | `Google/veo-3.1-fast`                   | text\_to\_video                                     | Google Veo 3.1 Fast - Fast and economical text-to-video generation. Supports up to 8 seconds at 720p/1080p.… |
| **Veo-3.1-Fast-I2V**                 | `Google/veo-3.1-fast-i2v`               | image\_to\_video                                    | Google Veo 3.1 Fast Image-to-Video - Generate video from an input image. Supports up to 8 seconds at…        |
| **Veo-3.1**                          | `Google/veo-3.1`                        | text\_to\_video                                     | Google Veo 3.1 Standard - Balanced quality and speed for text-to-video generation. Supports up to 8 seconds… |
| **Veo-3.1-I2V**                      | `Google/veo-3.1-i2v`                    | image\_to\_video                                    | Google Veo 3.1 Standard Image-to-Video - Generate high-quality video from an input image.                    |
| **Seedance-1.0-Pro-Image-to-Video**  | `Byteplus/seedance-1-0-pro-250528`      | image\_to\_video                                    | Pro-tier model, turns images into cinematic 1080p videos with smooth motion                                  |
| **Seedance-1.0-Pro-Text-to-Video**   | `Byteplus/seedance-1-0-pro-250528`      | text\_to\_video                                     | Pro-tier text-to-video, generates multi-shot 1080p videos with narrative flow                                |
| **Seedance-1.0-Lite-Image-to-Video** | `Byteplus/seedance-1-0-lite-i2v-250428` | image\_to\_video                                    | Lite version, quick image-to-video at up to 1080p, shorter sequences                                         |
| **Seedance-1.0-Lite-Text-to-Video**  | `Byteplus/seedance-1-0-lite-t2v-250428` | text\_to\_video                                     | Lite text-to-video, faster, lower compute, shorter outputs                                                   |

## Embeddings

Called through [Embeddings](/api-reference/inference-api/embeddings.md). See [Embeddings](/products/serverless-inference/embeddings.md).

| Model                       | Model ID                  | Parameters | Description                                                             |
| --------------------------- | ------------------------- | ---------- | ----------------------------------------------------------------------- |
| **Qwen3-Embedding-8B** 🇨🇦 | `Qwen/Qwen3-Embedding-8B` | 8B         | 8B embedding model, strong multilingual & code support, top MTEB scorer |

## Reranking

Called through [Rerank](/api-reference/inference-api/rerank.md). See [Reranking](/products/serverless-inference/reranking.md).

| Model                       | Model ID                  | Parameters | Description                                                                |
| --------------------------- | ------------------------- | ---------- | -------------------------------------------------------------------------- |
| **BGE-reranker-v2-m3** 🇨🇦 | `BAAI/bge-reranker-v2-m3` | 568M       | Multilingual reranker, query+passage → relevance score, lightweight & fast |

## Model access by tier

Every model carries a per-tier daily request cap, and on **Tier 1 that cap is 0 for most of the catalog** — a $5 deposit (Tier 2) is what unlocks it. Check the exact cap for your account under [Limits](https://console.nebulablock.com/limits) in the console, and see [Tiers and Rate Limits](/account/tiers-and-limits.md) for the full picture.

## See also

* [Serverless Inference overview](/products/serverless-inference.md)
* [List Models API](/api-reference/inference-api/list-models.md)
* [Tiers and Rate Limits](/account/tiers-and-limits.md)
