AIMs Catalog

AIMs Catalog#

The AIM catalog includes ready-to-deploy inference microservices for popular open-weight models, from lightweight instruction-tuned models to large mixture-of-experts systems. The catalog is constantly expanded to add new models and hardware-optimized profiles. To deploy models currently not available in the catalog but available through Hugging Face and supported by vLLM, utilize the AIM base container together with a custom profile (see here).

AIMs are available for AMD Instinct™, Radeon™, and EPYC™. Browse the catalog below to find a container for your model and hardware.

Model

Organization

Description

Resources

CohereLabs/command-a-reasoning-08-2025 (stable)

Cohere Labs

111B parameter language model with configurable reasoning and tool use capabilities.

Spec Deploy

deepseek-ai/DeepSeek-R1 (stable)

DeepSeek

671B parameter MoE reasoning model with 37B active parameters.

Spec Deploy

deepseek-ai/DeepSeek-R1-0528 (stable)

DeepSeek

671B parameter MoE reasoning model with 37B active parameters, updated version of DeepSeek-R1.

Spec Deploy

deepseek-ai/DeepSeek-V3.1 (stable)

DeepSeek

671B parameter MoE model with 37B active parameters supporting thinking and non-thinking modes.

Spec Deploy

deepseek-ai/DeepSeek-V3.1-Terminus (stable)

DeepSeek

671B parameter MoE model with 37B active parameters, refined for language consistency and agent tasks.

Spec Deploy

google/gemma-3-1b-it (stable)

Google

Gemma 3 1B IT is a lightweight instruction-tuned model supporting text generation with a 32K context window.

Spec Deploy

google/gemma-3-27b-it (stable)

Google

Gemma 3 27B IT is a multimodal instruction-tuned model supporting text and image input with a 128K context window.

Spec Deploy

google/gemma-4-31B-it (preview)

Google

Gemma 4 31B IT is a multimodal instruction-tuned model with text and image input, 256K native context, and Gemma 4 reasoning + tool-call parsers.

Spec Deploy

google/medgemma-27b-it (stable)

Google

Gemma 3-based 27B multimodal model fine-tuned for medical text and image tasks (X-ray, dermatology, ophthalmology, pathology, radiology reports).

Spec Deploy

meta-llama/Llama-3.1-405B-Instruct (stable)

Meta

Multilingual 405B parameter instruction-tuned language model for dialogue use cases.

Spec Deploy

meta-llama/Llama-3.1-8B-Instruct (stable)

Meta

Multilingual 8B parameter instruction-tuned language model for dialogue use cases.

Spec Deploy

meta-llama/Llama-3.2-1B-Instruct (stable)

Meta

Multilingual 1B parameter instruction-tuned language model for dialogue and on-device use cases.

Spec Deploy

meta-llama/Llama-3.2-3B-Instruct (stable)

Meta

Multilingual 3B parameter instruction-tuned language model for dialogue and on-device use cases.

Spec Deploy

meta-llama/Llama-3.3-70B-Instruct (stable)

Meta

Multilingual 70B parameter instruction-tuned language model for dialogue use cases.

Spec Deploy

MiniMaxAI/MiniMax-M2.5 (stable)

MiniMax

228B parameter mixture-of-experts language model with reasoning, tool calling, and coding capabilities.

Spec Deploy

MiniMaxAI/MiniMax-M3 (stable)

MiniMax

Native multimodal MoE model with 428B parameters, 23B active parameters, a 1M-token context window, reasoning, tool use, and coding capabilities.

Spec Deploy

mistralai/Ministral-3-14B-Instruct-2512 (stable)

Mistral AI

14B parameter instruction-tuned language model with vision and function calling capabilities.

Spec Deploy

mistralai/Ministral-3-14B-Reasoning-2512 (stable)

Mistral AI

14B parameter instruction-tuned language model with vision and function calling capabilities.

Spec Deploy

mistralai/Mistral-Large-3-675B-Instruct-2512 (stable)

Mistral AI

675B parameter granular MoE multimodal model with 41B active parameters and vision capabilities.

Spec Deploy

mistralai/Mistral-Small-24B-Instruct-2501 (stable)

Mistral AI

24B parameter instruction-tuned language model (Mistral Small 3) with native function calling. Text-only.

Spec Deploy

mistralai/Mistral-Small-3.2-24B-Instruct-2506 (stable)

Mistral AI

24B parameter instruction-tuned language model with vision and function calling capabilities.

Spec Deploy

mistralai/Mixtral-8x22B-Instruct-v0.1 (stable)

Mistral AI

Sparse MoE language model with 141B total parameters across 8 experts and function calling support.

Spec Deploy

mistralai/Mixtral-8x7B-Instruct-v0.1 (stable)

Mistral AI

Sparse MoE language model with 47B total parameters across 8 experts.

Spec Deploy

openai/gpt-oss-120b (stable)

OpenAI

Open-weight 117B parameter MoE model with 5.1B active parameters and configurable reasoning.

Spec Deploy

openai/gpt-oss-20b (stable)

OpenAI

Open-weight 21B parameter MoE model with 3.6B active parameters for lower-latency use cases.

Spec Deploy

Qwen/Qwen3-235B-A22B (stable)

Qwen

235B parameter MoE language model with 22B active parameters and dual thinking modes.

Spec Deploy

Qwen/Qwen3-32B (stable)

Qwen

32.8B parameter dense language model with dual thinking modes and multilingual support.

Spec Deploy

Qwen/Qwen3-Coder-Next (stable)

Qwen

80B parameter MoE coding agent model with 3B active parameters and hybrid attention architecture.

Spec Deploy

Qwen/Qwen3-VL-235B-A22B-Instruct (stable)

Qwen

236B parameter MoE vision-language model with 22B active parameters and multimodal capabilities.

Spec Deploy

Qwen/Qwen3-VL-235B-A22B-Thinking (stable)

Qwen

236B parameter MoE vision-language model with reasoning-enhanced thinking capabilities.

Spec Deploy

zai-org/GLM-4.7 (stable)

Z.ai

GLM-4.7 is a large language model with multi-turn conversation, tool use, and reasoning capabilities.

Spec Deploy

zai-org/GLM-5.2 (stable)

Z.ai

GLM-5.2 is a large Mixture-of-Experts language model with multi-turn conversation, tool use, and reasoning capabilities.

Spec Deploy

Model

Organization

Description

Resources

deepseek-ai/DeepSeek-R1-Distill-Qwen-14B (preview)

DeepSeek

14B parameter distilled reasoning model based on Qwen-2.5, optimized for CPU inference on AMD EPYC.

Spec Deploy

deepseek-ai/DeepSeek-R1-Distill-Qwen-7B (stable)

DeepSeek

7B parameter distilled reasoning model based on Qwen-2.5, optimized for CPU inference on AMD EPYC.

Spec Deploy

google/gemma-3-12b-it (stable)

Google

Gemma 3 12B Instruct multimodal model from Google DeepMind.

Spec Deploy

google/gemma-3-1b-it (stable)

Google

Gemma 3 1B Instruct multimodal model from Google DeepMind.

Spec Deploy

google/gemma-3-4b-it (stable)

Google

Gemma 3 4B Instruct multimodal model from Google DeepMind.

Spec Deploy

google/gemma-4-E4B-it (stable)

Google

Gemma 4 E4B is Google’s efficient instruction-tuned model with a ~4B effective-parameter MatFormer architecture for low-latency on-device and CPU inference.

Spec Deploy

meta-llama/Llama-3.1-8B-Instruct (stable)

Meta

Multilingual 8B parameter instruction-tuned language model for dialogue use cases.

Spec Deploy

meta-llama/Llama-3.2-1B-Instruct (stable)

Meta

Multilingual 1B parameter instruction-tuned language model for dialogue and on-device use cases.

Spec Deploy

meta-llama/Llama-3.2-3B-Instruct (stable)

Meta

Multilingual 3B parameter instruction-tuned language model for dialogue and on-device use cases.

Spec Deploy

Qwen/Qwen2.5-Coder-7B-Instruct (stable)

Qwen

7B parameter coding and instruction-following model optimized for CPU inference on AMD EPYC.

Spec Deploy

Qwen/Qwen2.5-VL-7B-Instruct (preview)

Qwen

7B parameter vision-language model optimized for CPU inference on AMD EPYC.

Spec Deploy

Qwen/Qwen3-0.6B (stable)

Qwen

Reasoning-enhanced 0.6B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations.

Spec Deploy

Qwen/Qwen3-1.7B (stable)

Qwen

Reasoning-enhanced 1.7B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations.

Spec Deploy

Qwen/Qwen3-30B-A3B (stable)

Qwen

Mixture-of-experts 30B (3B active) LLM with thinking/non-thinking mode switching, advanced reasoning, agent capabilities, and 100+ language support.

Spec Deploy

Qwen/Qwen3-4B (stable)

Qwen

Reasoning-enhanced 4B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations.

Spec Deploy

Qwen/Qwen3-8B (stable)

Qwen

Reasoning-enhanced 8B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations.

Spec Deploy

Qwen/Qwen3.5-4B (stable)

Qwen

Qwen3.5-4B is a 4B parameter LLM with thinking/non-thinking dual-mode reasoning, strong math and coding ability, and multilingual support.

Spec Deploy

Qwen/Qwen3.5-9B (stable)

Qwen

Qwen3.5-9B is a 9B parameter LLM with thinking/non-thinking dual-mode reasoning, strong math and coding ability, and multilingual support.

Spec Deploy

Qwen/Qwen3.6-35B-A3B (stable)

Qwen

Qwen3.6-35B-A3B is a Mixture-of-Experts LLM with 35B total parameters and ~3B active per token, balancing high quality with efficient inference.

Spec Deploy

ibm-granite/granite-3.3-2b-instruct (stable)

ibm-granite

2B parameter instruction-tuned language model with 128K context length and improved reasoning capabilities.

Spec Deploy

ibm-granite/granite-4.1-3b (stable)

ibm-granite

3B parameter long-context instruct model with 131K context length and enhanced capabilities for summarization, RAG, and code tasks.

Spec Deploy

ibm-granite/granite-4.1-8b (stable)

ibm-granite

8B parameter long-context instruct model with 131K context length and enhanced capabilities for summarization, RAG, and code tasks.

Spec Deploy

microsoft/Phi-3.5-mini-instruct (stable)

microsoft

3.8B parameter language model optimized for reasoning, math, and code tasks in memory-constrained environments.

Spec Deploy

unsloth/gpt-oss-20b-BF16 (stable)

unsloth

gpt-oss-20B is OpenAI’s open-weight 20B Mixture-of-Experts model (BF16 conversion by Unsloth) with strong reasoning and tool-use capabilities.

Spec Deploy

Model

Organization

Description

Resources

google/gemma-3n-E4B-it (preview)

Google

Gemma 3n E4B IT is a gated multimodal instruction-tuned model supporting text, image, video, and audio inputs.

Spec Deploy

meta-llama/Llama-3.1-8B-Instruct (preview)

Meta

Multilingual 8B parameter instruction-tuned language model for dialogue use cases.

Spec Deploy

Qwen/Qwen3-VL-8B-Instruct (preview)

Qwen

8B parameter vision-language model with advanced multimodal reasoning.

Spec Deploy

Qwen/Qwen3.5-9B (preview)

Qwen

9B parameter hybrid language model with Gated DeltaNet and dual thinking modes.

Spec Deploy

zai-org/GLM-4.7-Flash (preview)

Z.ai

30B-A3B MoE language model with balanced performance and efficiency for lightweight deployment.

Spec Deploy