AIMs Catalog#
The AIM catalog includes ready-to-deploy inference microservices for popular open-weight models, from lightweight instruction-tuned models to large mixture-of-experts systems. The catalog is constantly expanded to add new models and hardware-optimized profiles. To deploy models currently not available in the catalog but available through Hugging Face and supported by vLLM, utilize the AIM base container together with a custom profile (see here).
AIMs are available for AMD Instinct™, Radeon™, and EPYC™. Browse the catalog below to find a container for your model and hardware.
Model |
Organization |
Description |
Resources |
|---|---|---|---|
Cohere Labs |
111B parameter language model with configurable reasoning and tool use capabilities. |
||
deepseek-ai/DeepSeek-R1 (stable) |
DeepSeek |
671B parameter MoE reasoning model with 37B active parameters. |
|
deepseek-ai/DeepSeek-R1-0528 (stable) |
DeepSeek |
671B parameter MoE reasoning model with 37B active parameters, updated version of DeepSeek-R1. |
|
deepseek-ai/DeepSeek-V3.1 (stable) |
DeepSeek |
671B parameter MoE model with 37B active parameters supporting thinking and non-thinking modes. |
|
deepseek-ai/DeepSeek-V3.1-Terminus (stable) |
DeepSeek |
671B parameter MoE model with 37B active parameters, refined for language consistency and agent tasks. |
|
google/gemma-3-1b-it (stable) |
Gemma 3 1B IT is a lightweight instruction-tuned model supporting text generation with a 32K context window. |
||
google/gemma-3-27b-it (stable) |
Gemma 3 27B IT is a multimodal instruction-tuned model supporting text and image input with a 128K context window. |
||
google/gemma-4-31B-it (preview) |
Gemma 4 31B IT is a multimodal instruction-tuned model with text and image input, 256K native context, and Gemma 4 reasoning + tool-call parsers. |
||
google/medgemma-27b-it (stable) |
Gemma 3-based 27B multimodal model fine-tuned for medical text and image tasks (X-ray, dermatology, ophthalmology, pathology, radiology reports). |
||
meta-llama/Llama-3.1-405B-Instruct (stable) |
Meta |
Multilingual 405B parameter instruction-tuned language model for dialogue use cases. |
|
meta-llama/Llama-3.1-8B-Instruct (stable) |
Meta |
Multilingual 8B parameter instruction-tuned language model for dialogue use cases. |
|
meta-llama/Llama-3.2-1B-Instruct (stable) |
Meta |
Multilingual 1B parameter instruction-tuned language model for dialogue and on-device use cases. |
|
meta-llama/Llama-3.2-3B-Instruct (stable) |
Meta |
Multilingual 3B parameter instruction-tuned language model for dialogue and on-device use cases. |
|
meta-llama/Llama-3.3-70B-Instruct (stable) |
Meta |
Multilingual 70B parameter instruction-tuned language model for dialogue use cases. |
|
MiniMaxAI/MiniMax-M2.5 (stable) |
MiniMax |
228B parameter mixture-of-experts language model with reasoning, tool calling, and coding capabilities. |
|
Mistral AI |
14B parameter instruction-tuned language model with vision and function calling capabilities. |
||
Mistral AI |
14B parameter instruction-tuned language model with vision and function calling capabilities. |
||
Mistral AI |
675B parameter granular MoE multimodal model with 41B active parameters and vision capabilities. |
||
Mistral AI |
24B parameter instruction-tuned language model (Mistral Small 3) with native function calling. Text-only. |
||
Mistral AI |
24B parameter instruction-tuned language model with vision and function calling capabilities. |
||
Mistral AI |
Sparse MoE language model with 141B total parameters across 8 experts and function calling support. |
||
mistralai/Mixtral-8x7B-Instruct-v0.1 (stable) |
Mistral AI |
Sparse MoE language model with 47B total parameters across 8 experts. |
|
openai/gpt-oss-120b (stable) |
OpenAI |
Open-weight 117B parameter MoE model with 5.1B active parameters and configurable reasoning. |
|
openai/gpt-oss-20b (stable) |
OpenAI |
Open-weight 21B parameter MoE model with 3.6B active parameters for lower-latency use cases. |
|
Qwen/Qwen3-235B-A22B (stable) |
Qwen |
235B parameter MoE language model with 22B active parameters and dual thinking modes. |
|
Qwen/Qwen3-32B (stable) |
Qwen |
32.8B parameter dense language model with dual thinking modes and multilingual support. |
|
Qwen/Qwen3-Coder-Next (stable) |
Qwen |
80B parameter MoE coding agent model with 3B active parameters and hybrid attention architecture. |
|
Qwen/Qwen3-VL-235B-A22B-Instruct (stable) |
Qwen |
236B parameter MoE vision-language model with 22B active parameters and multimodal capabilities. |
|
Qwen/Qwen3-VL-235B-A22B-Thinking (stable) |
Qwen |
236B parameter MoE vision-language model with reasoning-enhanced thinking capabilities. |
|
zai-org/GLM-4.7 (stable) |
Z.ai |
GLM-4.7 is a large language model with multi-turn conversation, tool use, and reasoning capabilities. |
Model |
Organization |
Description |
Resources |
|---|---|---|---|
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B (preview) |
DeepSeek |
14B parameter distilled reasoning model based on Qwen-2.5, optimized for CPU inference on AMD EPYC. |
|
DeepSeek |
7B parameter distilled reasoning model based on Qwen-2.5, optimized for CPU inference on AMD EPYC. |
||
google/gemma-3-12b-it (stable) |
Gemma 3 12B Instruct multimodal model from Google DeepMind. |
||
google/gemma-3-1b-it (stable) |
Gemma 3 1B Instruct multimodal model from Google DeepMind. |
||
google/gemma-3-4b-it (stable) |
Gemma 3 4B Instruct multimodal model from Google DeepMind. |
||
google/gemma-4-E4B-it (stable) |
Gemma 4 E4B is Google’s efficient instruction-tuned model with a ~4B effective-parameter MatFormer architecture for low-latency on-device and CPU inference. |
||
meta-llama/Llama-3.1-8B-Instruct (stable) |
Meta |
Multilingual 8B parameter instruction-tuned language model for dialogue use cases. |
|
meta-llama/Llama-3.2-1B-Instruct (stable) |
Meta |
Multilingual 1B parameter instruction-tuned language model for dialogue and on-device use cases. |
|
meta-llama/Llama-3.2-3B-Instruct (stable) |
Meta |
Multilingual 3B parameter instruction-tuned language model for dialogue and on-device use cases. |
|
Qwen/Qwen2.5-Coder-7B-Instruct (stable) |
Qwen |
7B parameter coding and instruction-following model optimized for CPU inference on AMD EPYC. |
|
Qwen/Qwen2.5-VL-7B-Instruct (preview) |
Qwen |
7B parameter vision-language model optimized for CPU inference on AMD EPYC. |
|
Qwen/Qwen3-0.6B (stable) |
Qwen |
Reasoning-enhanced 0.6B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations. |
|
Qwen/Qwen3-1.7B (stable) |
Qwen |
Reasoning-enhanced 1.7B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations. |
|
Qwen/Qwen3-30B-A3B (stable) |
Qwen |
Mixture-of-experts 30B (3B active) LLM with thinking/non-thinking mode switching, advanced reasoning, agent capabilities, and 100+ language support. |
|
Qwen/Qwen3-4B (stable) |
Qwen |
Reasoning-enhanced 4B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations. |
|
Qwen/Qwen3-8B (stable) |
Qwen |
Reasoning-enhanced 8B parameter LLM with thinking/non-thinking mode switching, excelling in math, coding, and multi-turn conversations. |
|
Qwen/Qwen3.5-4B (stable) |
Qwen |
Qwen3.5-4B is a 4B parameter LLM with thinking/non-thinking dual-mode reasoning, strong math and coding ability, and multilingual support. |
|
Qwen/Qwen3.5-9B (stable) |
Qwen |
Qwen3.5-9B is a 9B parameter LLM with thinking/non-thinking dual-mode reasoning, strong math and coding ability, and multilingual support. |
|
Qwen/Qwen3.6-35B-A3B (stable) |
Qwen |
Qwen3.6-35B-A3B is a Mixture-of-Experts LLM with 35B total parameters and ~3B active per token, balancing high quality with efficient inference. |
|
ibm-granite/granite-3.3-2b-instruct (stable) |
ibm-granite |
2B parameter instruction-tuned language model with 128K context length and improved reasoning capabilities. |
|
ibm-granite/granite-4.1-3b (stable) |
ibm-granite |
3B parameter long-context instruct model with 131K context length and enhanced capabilities for summarization, RAG, and code tasks. |
|
ibm-granite/granite-4.1-8b (stable) |
ibm-granite |
8B parameter long-context instruct model with 131K context length and enhanced capabilities for summarization, RAG, and code tasks. |
|
microsoft/Phi-3.5-mini-instruct (stable) |
microsoft |
3.8B parameter language model optimized for reasoning, math, and code tasks in memory-constrained environments. |
|
unsloth/gpt-oss-20b-BF16 (stable) |
unsloth |
gpt-oss-20B is OpenAI’s open-weight 20B Mixture-of-Experts model (BF16 conversion by Unsloth) with strong reasoning and tool-use capabilities. |
Model |
Organization |
Description |
Resources |
|---|---|---|---|
google/gemma-3n-E4B-it (preview) |
Gemma 3n E4B IT is a gated multimodal instruction-tuned model supporting text, image, video, and audio inputs. |
||
meta-llama/Llama-3.1-8B-Instruct (preview) |
Meta |
Multilingual 8B parameter instruction-tuned language model for dialogue use cases. |
|
Qwen/Qwen3-VL-8B-Instruct (preview) |
Qwen |
8B parameter vision-language model with advanced multimodal reasoning. |
|
Qwen/Qwen3.5-9B (preview) |
Qwen |
9B parameter hybrid language model with Gated DeltaNet and dual thinking modes. |
|
zai-org/GLM-4.7-Flash (preview) |
Z.ai |
30B-A3B MoE language model with balanced performance and efficiency for lightweight deployment. |