Granite Models Find Granite models that fit your available resources and offer the required capabilities.Learn more
Compatible GPUs for these 27 models: CPU or entry-level GPU (T4, RTX 3050+), NVIDIA T4, RTX 3060/4060, NVIDIA RTX 4080, L4, A10, NVIDIA RTX 4070, RTX 3080
Granite-Speech-5.0-470M-TurboCTC
Resource Constraints
Est. VRAM: 2GB at BF16
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Transcription (ASR)
Model Details
Focus sentinel Model Family Speech
Parameters 470M Context Window N/A License Apache 2.0 Resource Constraints Est. VRAM: 2GB at BF16 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Batch transcription High-throughput pipelines Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 0.88GB 2GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-speech-5.0-470m-turboctc at revision 18ca3c1d on 2026-09-03. Run Granite Serving instructions for Granite-Speech-5.0-470M-TurboCTC aren't published here. See the ibm-granite/granite-speech-5.0-470m-turboctc model card on Hugging Face for how to run it.
When to Choose Maximum throughput batch transcription where speed is the priority; trades some accuracy for faster non-autoregressive decoding.
Focus sentinel Granite-Speech-5.0-470M-TurboCTC NC
Resource Constraints
Est. VRAM: 2GB at BF16
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Transcription (ASR)
Model Details
Focus sentinel Model Family Speech
Parameters 470M Context Window N/A License CC-BY-NC-SA-4.0 Resource Constraints Est. VRAM: 2GB at BF16 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Batch transcription High-throughput pipelines Noncommercial use only Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 0.88GB 2GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-speech-5.0-470m-turboctc-nc at revision 0eb7b4fe on 2026-09-03. Run Granite Serving instructions for Granite-Speech-5.0-470M-TurboCTC NC aren't published here. See the ibm-granite/granite-speech-5.0-470m-turboctc-nc model card on Hugging Face for how to run it.
When to Choose Maximum throughput batch transcription where speed is the priority; trades some accuracy for faster non-autoregressive decoding.
Focus sentinel Resource Constraints
Est. VRAM: 8GB at BF16, down to 3GB at Q2_K
Runs on CPU or an entry-level GPU
Supported GPUs: NVIDIA T4, RTX 3060/4060
Capabilities: Document Summarization, Multilingual Generation, Instruction Following, Tool-calling, Agentic Workflows, Coding, Math
Model Details
Focus sentinel Model Family Language
Parameters 3B Context Window 128K License Apache 2.0 Resource Constraints Est. VRAM: 8GB at BF16, down to 3GB at Q2_K Runs on CPU or an entry-level GPU Supported GPUs: NVIDIA T4, RTX 3060/4060 Capabilities AI assistants RAG pipelines Tool-calling Agentic workflows Coding Multilingual generation Structured JSON output Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 6.82GB 8GB Consumer GPU (5-12GB VRAM) FP8 safetensors 3.89GB 5GB Consumer GPU (5-12GB VRAM) Q8_0 GGUF 3.63GB 5GB Consumer GPU (5-12GB VRAM) Q6_K GGUF 2.8GB 4GB CPU-friendly (up to 4GB VRAM) NVFP4 safetensors 2.61GB 4GB CPU-friendly (up to 4GB VRAM) Q5_1 GGUF 2.58GB 4GB CPU-friendly (up to 4GB VRAM) MXFP4 safetensors 2.51GB 4GB CPU-friendly (up to 4GB VRAM) Q5_K_M GGUF 2.43GB 4GB CPU-friendly (up to 4GB VRAM) Q5_0 GGUF 2.38GB 4GB CPU-friendly (up to 4GB VRAM) Q5_K_S GGUF 2.38GB 4GB CPU-friendly (up to 4GB VRAM) Q4_1 GGUF 2.18GB 3GB CPU-friendly (up to 4GB VRAM) Q4_K_M Recommended
GGUF 2.09GB 3GB CPU-friendly (up to 4GB VRAM) Q4_K_S GGUF 2GB 3GB CPU-friendly (up to 4GB VRAM) Q4_0 GGUF 1.98GB 3GB CPU-friendly (up to 4GB VRAM) Q3_K_L Heavily quantized
GGUF 1.84GB 3GB CPU-friendly (up to 4GB VRAM) Q3_K_M Heavily quantized
GGUF 1.71GB 3GB CPU-friendly (up to 4GB VRAM) Q3_K_S Heavily quantized
GGUF 1.56GB 3GB CPU-friendly (up to 4GB VRAM) Q2_K Heavily quantized
GGUF 1.36GB 3GB CPU-friendly (up to 4GB VRAM)
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-4.2-3b at revision e459acce on 2026-09-03. Run Granite Run Granite 4.2 3B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-4.2-3b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-4.2-3b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Tasks requiring deeper reasoning or multi-step agentic chains where output quality is the top priority and compute is available.
Focus sentinel Resource Constraints
Est. VRAM: 19GB at BF16, down to 5GB at Q2_K
Supported GPUs: NVIDIA RTX 4090, L4, A10G
Capabilities: Document Summarization, Multilingual Generation, Instruction Following, Tool-calling, Agentic Workflows, Coding, Math
Model Details
Focus sentinel Model Family Language
Parameters 8B Context Window 128K License Apache 2.0 Resource Constraints Est. VRAM: 19GB at BF16, down to 5GB at Q2_K Supported GPUs: NVIDIA RTX 4090, L4, A10G Capabilities AI assistants RAG pipelines Tool-calling Agentic workflows Coding Multilingual generation Structured JSON output Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 16.38GB 19GB Prosumer GPU (13-24GB VRAM) FP8 safetensors 8.96GB 11GB Consumer GPU (5-12GB VRAM) Q8_0 GGUF 8.7GB 11GB Consumer GPU (5-12GB VRAM) Q6_K GGUF 6.72GB 9GB Consumer GPU (5-12GB VRAM) Q5_1 GGUF 6.17GB 8GB Consumer GPU (5-12GB VRAM) Q5_K_M GGUF 5.82GB 8GB Consumer GPU (5-12GB VRAM) NVFP4 safetensors 5.71GB 7GB Consumer GPU (5-12GB VRAM) Q5_0 GGUF 5.68GB 7GB Consumer GPU (5-12GB VRAM) Q5_K_S GGUF 5.68GB 7GB Consumer GPU (5-12GB VRAM) MXFP4 safetensors 5.47GB 7GB Consumer GPU (5-12GB VRAM) Q4_1 GGUF 5.2GB 7GB Consumer GPU (5-12GB VRAM) Q4_K_M Recommended
GGUF 4.98GB 7GB Consumer GPU (5-12GB VRAM) Q4_K_S GGUF 4.74GB 6GB Consumer GPU (5-12GB VRAM) Q4_0 GGUF 4.71GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_L Heavily quantized
GGUF 4.38GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_M Heavily quantized
GGUF 4.05GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_S Heavily quantized
GGUF 3.67GB 5GB Consumer GPU (5-12GB VRAM) Q2_K Heavily quantized
GGUF 3.18GB 5GB Consumer GPU (5-12GB VRAM)
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-4.2-8b at revision 7fce579a on 2026-09-03. Run Granite Run Granite 4.2 8B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-4.2-8b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-4.2-8b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Tasks requiring deeper reasoning or multi-step agentic chains where output quality is the top priority and compute is available.
Focus sentinel Resource Constraints
Est. VRAM: 61GB at BF16, down to 13GB at Q2_K
Supported GPUs: NVIDIA A100 (80GB), H100
Capabilities: Document Summarization, Multilingual Generation, Instruction Following, Tool-calling, Agentic Workflows, Coding, Math
Model Details
Focus sentinel Model Family Language
Parameters 30B Context Window 128K License Apache 2.0 Resource Constraints Est. VRAM: 61GB at BF16, down to 13GB at Q2_K Supported GPUs: NVIDIA A100 (80GB), H100 Capabilities AI assistants RAG pipelines Tool-calling Agentic workflows Coding Multilingual generation Structured JSON output Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 54.53GB 61GB Datacenter GPU (25-80GB VRAM) Q8_0 GGUF 28.98GB 33GB Datacenter GPU (25-80GB VRAM) FP8 safetensors 28.04GB 32GB Datacenter GPU (25-80GB VRAM) Q6_K GGUF 22.37GB 26GB Datacenter GPU (25-80GB VRAM) Q5_1 GGUF 20.48GB 24GB Prosumer GPU (13-24GB VRAM) Q5_K_M GGUF 19.35GB 23GB Prosumer GPU (13-24GB VRAM) Q5_0 GGUF 18.8GB 22GB Prosumer GPU (13-24GB VRAM) Q5_K_S GGUF 18.8GB 22GB Prosumer GPU (13-24GB VRAM) Q4_1 GGUF 17.12GB 20GB Prosumer GPU (13-24GB VRAM) Q4_K_M Recommended
GGUF 16.5GB 20GB Prosumer GPU (13-24GB VRAM) NVFP4 safetensors 16.44GB 20GB Prosumer GPU (13-24GB VRAM) MXFP4 safetensors 15.61GB 19GB Prosumer GPU (13-24GB VRAM) Q4_K_S GGUF 15.57GB 19GB Prosumer GPU (13-24GB VRAM) Q4_0 GGUF 15.44GB 18GB Prosumer GPU (13-24GB VRAM) Q3_K_L Heavily quantized
GGUF 14.26GB 17GB Prosumer GPU (13-24GB VRAM) Q3_K_M Heavily quantized
GGUF 13.16GB 16GB Prosumer GPU (13-24GB VRAM) Q3_K_S Heavily quantized
GGUF 11.87GB 15GB Prosumer GPU (13-24GB VRAM) Q2_K Heavily quantized
GGUF 10.11GB 13GB Prosumer GPU (13-24GB VRAM)
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-4.2-30b at revision 8b445a5c on 2026-09-03. Run Granite Run Granite 4.2 30B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-4.2-30b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-4.2-30b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Tasks requiring deeper reasoning or multi-step agentic chains where output quality is the top priority and compute is available.
Focus sentinel Resource Constraints
Est. VRAM: 8GB at BF16, down to 3GB at Q2_K
Runs on CPU or an entry-level GPU
Supported GPUs: NVIDIA T4, RTX 3060/4060
Capabilities: Document Summarization, Multilingual Generation, Instruction Following, Tool-calling
Model Details
Focus sentinel Model Family Language
Parameters 3B Context Window 128K License Apache 2.0 Resource Constraints Est. VRAM: 8GB at BF16, down to 3GB at Q2_K Runs on CPU or an entry-level GPU Supported GPUs: NVIDIA T4, RTX 3060/4060 Capabilities AI assistants RAG pipelines Tool-calling Agentic workflows Coding Multilingual generation Structured JSON output Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 6.34GB 8GB Consumer GPU (5-12GB VRAM) Q8_0 GGUF 3.37GB 5GB Consumer GPU (5-12GB VRAM) Q6_K GGUF 2.6GB 4GB CPU-friendly (up to 4GB VRAM) Q5_1 GGUF 2.4GB 4GB CPU-friendly (up to 4GB VRAM) Q5_K_M GGUF 2.27GB 4GB CPU-friendly (up to 4GB VRAM) Q5_0 GGUF 2.21GB 4GB CPU-friendly (up to 4GB VRAM) Q5_K_S GGUF 2.21GB 4GB CPU-friendly (up to 4GB VRAM) Q4_1 GGUF 2.03GB 3GB CPU-friendly (up to 4GB VRAM) Q4_K_M Recommended
GGUF 1.96GB 3GB CPU-friendly (up to 4GB VRAM) Q4_K_S GGUF 1.86GB 3GB CPU-friendly (up to 4GB VRAM) Q4_0 GGUF 1.85GB 3GB CPU-friendly (up to 4GB VRAM) Q3_K_L Heavily quantized
GGUF 1.74GB 3GB CPU-friendly (up to 4GB VRAM) Q3_K_M Heavily quantized
GGUF 1.61GB 3GB CPU-friendly (up to 4GB VRAM) Q3_K_S Heavily quantized
GGUF 1.46GB 3GB CPU-friendly (up to 4GB VRAM) Q2_K Heavily quantized
GGUF 1.28GB 3GB CPU-friendly (up to 4GB VRAM)
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-4.1-3b at revision c0650403 on 2026-09-03. Run Granite Run Granite 4.1 3B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-4.1-3b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-4.1-3b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Resource-constrained environments; on-device or real-time inference; cost matters more than peak accuracy.
Focus sentinel Resource Constraints
Est. VRAM: 19GB at BF16, down to 5GB at Q2_K
Supported GPUs: NVIDIA RTX 4090, L4, A10G
Capabilities: Document Summarization, Multilingual Generation, Instruction Following, Tool-calling, Agentic Workflows
Model Details
Focus sentinel Model Family Language
Parameters 8B Context Window 128K License Apache 2.0 Resource Constraints Est. VRAM: 19GB at BF16, down to 5GB at Q2_K Supported GPUs: NVIDIA RTX 4090, L4, A10G Capabilities AI assistants RAG pipelines Tool-calling Agentic workflows Coding Multilingual generation Structured JSON output Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 16.38GB 19GB Prosumer GPU (13-24GB VRAM) Q8_0 GGUF 8.7GB 11GB Consumer GPU (5-12GB VRAM) Q6_K GGUF 6.72GB 9GB Consumer GPU (5-12GB VRAM) Q5_1 GGUF 6.17GB 8GB Consumer GPU (5-12GB VRAM) Q5_K_M GGUF 5.82GB 8GB Consumer GPU (5-12GB VRAM) Q5_0 GGUF 5.68GB 7GB Consumer GPU (5-12GB VRAM) Q5_K_S GGUF 5.68GB 7GB Consumer GPU (5-12GB VRAM) Q4_1 GGUF 5.2GB 7GB Consumer GPU (5-12GB VRAM) Q4_K_M Recommended
GGUF 4.98GB 7GB Consumer GPU (5-12GB VRAM) Q4_K_S GGUF 4.74GB 6GB Consumer GPU (5-12GB VRAM) Q4_0 GGUF 4.71GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_L Heavily quantized
GGUF 4.38GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_M Heavily quantized
GGUF 4.05GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_S Heavily quantized
GGUF 3.67GB 5GB Consumer GPU (5-12GB VRAM) Q2_K Heavily quantized
GGUF 3.18GB 5GB Consumer GPU (5-12GB VRAM)
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-4.1-8b at revision 1504002f on 2026-09-03. Run Granite Run Granite 4.1 8B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-4.1-8b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-4.1-8b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Best all-around choice for most production workloads — strong across coding, RAG, and instruction following without the resource cost of 30B.
Focus sentinel Resource Constraints
Est. VRAM: 61GB at BF16, down to 12GB at Q2_K
Supported GPUs: NVIDIA A100 (80GB), H100
Capabilities: Document Summarization, Multilingual Generation, Instruction Following, Tool-calling, Agentic Workflows, Coding, Math
Model Details
Focus sentinel Model Family Language
Parameters 30B Context Window 128K License Apache 2.0 Resource Constraints Est. VRAM: 61GB at BF16, down to 12GB at Q2_K Supported GPUs: NVIDIA A100 (80GB), H100 Capabilities AI assistants RAG pipelines Tool-calling Agentic workflows Coding Multilingual generation Structured JSON output Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 53.77GB 61GB Datacenter GPU (25-80GB VRAM) Q8_0 GGUF 28.57GB 33GB Datacenter GPU (25-80GB VRAM) Q6_K GGUF 22.06GB 26GB Datacenter GPU (25-80GB VRAM) Q5_1 GGUF 20.19GB 24GB Prosumer GPU (13-24GB VRAM) Q5_K_M GGUF 19.09GB 22GB Prosumer GPU (13-24GB VRAM) Q5_0 GGUF 18.54GB 22GB Prosumer GPU (13-24GB VRAM) Q5_K_S GGUF 18.54GB 22GB Prosumer GPU (13-24GB VRAM) Q4_1 GGUF 16.88GB 20GB Prosumer GPU (13-24GB VRAM) Q4_K_M Recommended
GGUF 16.29GB 19GB Prosumer GPU (13-24GB VRAM) Q4_K_S GGUF 15.35GB 18GB Prosumer GPU (13-24GB VRAM) Q4_0 GGUF 15.23GB 18GB Prosumer GPU (13-24GB VRAM) Q3_K_L Heavily quantized
GGUF 14.09GB 17GB Prosumer GPU (13-24GB VRAM) Q3_K_M Heavily quantized
GGUF 13GB 16GB Prosumer GPU (13-24GB VRAM) Q3_K_S Heavily quantized
GGUF 11.71GB 14GB Prosumer GPU (13-24GB VRAM) Q2_K Heavily quantized
GGUF 9.99GB 12GB Consumer GPU (5-12GB VRAM)
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-4.1-30b at revision 4fae6278 on 2026-09-03. Run Granite Run Granite 4.1 30B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-4.1-30b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-4.1-30b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Tasks requiring deeper reasoning or multi-step agentic chains where output quality is the top priority and compute is available.
Focus sentinel Resource Constraints
Est. VRAM: 9GB at BF16, down to 4GB at Q4_K_M
Runs on CPU or an entry-level GPU
Supported GPUs: NVIDIA RTX 4070, RTX 3080, L4
Capabilities: Chart Extraction, Table Extraction, KVP Extraction
Model Details
Focus sentinel Model Family Vision
Parameters 4B Context Window N/A License Apache 2.0 Resource Constraints Est. VRAM: 9GB at BF16, down to 4GB at Q4_K_M Runs on CPU or an entry-level GPU Supported GPUs: NVIDIA RTX 4070, RTX 3080, L4 Capabilities Document extraction tasks Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 7.45GB 9GB Consumer GPU (5-12GB VRAM) Q8_0 GGUF 4.45GB * 6GB Consumer GPU (5-12GB VRAM) Q6_K GGUF 3.69GB * 5GB Consumer GPU (5-12GB VRAM) Q5_K_M GGUF 3.35GB * 5GB Consumer GPU (5-12GB VRAM) Q4_K_M Recommended
GGUF 3.04GB * 4GB CPU-friendly (up to 4GB VRAM)
* Includes the multimodal projector, which loads alongside the language weights.
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-vision-4.1-4b at revision 37d591f0 on 2026-09-03. Run Granite Run Granite Vision 4.1 4B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-vision-4.1-4b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-vision-4.1-4b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Latest and most capable vision model; preferred for all document and image understanding tasks.
Focus sentinel Resource Constraints
Est. VRAM: 6GB at BF16, down to 3GB at Q4_K_M
Runs on CPU or an entry-level GPU
Supported GPUs: NVIDIA T4, RTX 3060/4060
Capabilities: Transcription (ASR), Automatic Speech Translation (AST)
Model Details
Focus sentinel Model Family Speech
Parameters 2B Context Window N/A License Apache 2.0 Resource Constraints Est. VRAM: 6GB at BF16, down to 3GB at Q4_K_M Runs on CPU or an entry-level GPU Supported GPUs: NVIDIA T4, RTX 3060/4060 Capabilities ASR and AST that may require punctuation & casing and/or keyword-list biasing Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 4.31GB 6GB Consumer GPU (5-12GB VRAM) Q8_0 GGUF 2.9GB * 4GB CPU-friendly (up to 4GB VRAM) Q6_K GGUF 2.49GB * 4GB CPU-friendly (up to 4GB VRAM) Q5_K_M GGUF 2.31GB * 4GB CPU-friendly (up to 4GB VRAM) Q4_K_M Recommended
GGUF 2.14GB * 3GB CPU-friendly (up to 4GB VRAM)
* Includes the multimodal projector, which loads alongside the language weights.
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-speech-4.1-2b at revision de575db6 on 2026-09-03. Run Granite Run Granite Speech 4.1 2B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-speech-4.1-2b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-speech-4.1-2b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Standard choice for most ASR and transcription tasks across EN, FR, DE, ES, PT, JA.
Focus sentinel Granite Speech 4.1 2B Plus
Resource Constraints
Est. VRAM: 5GB at BF16, down to 3GB at Q4_K_M
Runs on CPU or an entry-level GPU
Supported GPUs: NVIDIA T4, RTX 3060/4060
Capabilities: Transcription (ASR), Automatic Speech Translation (AST), Speaker Diarization, Word-level Timestamps
Model Details
Focus sentinel Model Family Speech
Parameters 2B Context Window N/A License Apache 2.0 Resource Constraints Est. VRAM: 5GB at BF16, down to 3GB at Q4_K_M Runs on CPU or an entry-level GPU Supported GPUs: NVIDIA T4, RTX 3060/4060 Capabilities Speaker diarization Word-level timestamps Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 3.93GB 5GB Consumer GPU (5-12GB VRAM) Q8_0 GGUF 2.71GB * 4GB CPU-friendly (up to 4GB VRAM) Q6_K GGUF 2.34GB * 4GB CPU-friendly (up to 4GB VRAM) Q5_K_M GGUF 2.18GB * 3GB CPU-friendly (up to 4GB VRAM) Q4_K_M Recommended
GGUF 2.04GB * 3GB CPU-friendly (up to 4GB VRAM)
* Includes the multimodal projector, which loads alongside the language weights.
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-speech-4.1-2b-plus at revision 1454e6e1 on 2026-09-03. Run Granite Run Granite Speech 4.1 2B Plus with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-speech-4.1-2b-plus3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-speech-4.1-2b-plus", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose When you need to identify who said what or require precise word-level timestamps for downstream processing.
Focus sentinel Granite Speech 4.1 2B NAR
Resource Constraints
Est. VRAM: 5GB at BF16
Supported GPUs: NVIDIA T4, RTX 3060/4060
Capabilities: Transcription (ASR)
Model Details
Focus sentinel Model Family Speech
Parameters 2B Context Window N/A License apache-2.0 Resource Constraints Est. VRAM: 5GB at BF16 Supported GPUs: NVIDIA T4, RTX 3060/4060 Capabilities Batch transcription High-throughput pipelines Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 4.2GB 5GB Consumer GPU (5-12GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-speech-4.1-2b-nar at revision a1e3416e on 2026-09-03. Run Granite Run Granite Speech 4.1 2B NAR with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-speech-4.1-2b-nar3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-speech-4.1-2b-nar", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Maximum throughput batch transcription where speed is the priority; trades some accuracy for faster non-autoregressive decoding.
Focus sentinel Granite Embedding 97M Multilingual R2
Resource Constraints
Est. VRAM: 1GB at BF16
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Retrieval, Multilingual Retrieval, Code Retrieval
Model Details
Focus sentinel Model Family Embedding
Parameters 97M Context Window 32K License apache-2.0 Resource Constraints Est. VRAM: 1GB at BF16 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 0.18GB 1GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-embedding-97m-multilingual-r2 at revision 835ad140 on 2026-09-03. Run Granite Run Granite Embedding 97M Multilingual R2 with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-embedding-97m-multilingual-r23. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-embedding-97m-multilingual-r2", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Best sub-100M multilingual option with 32K context; ideal when speed and multi-language support are both required.
Focus sentinel Granite Embedding 311M Multilingual R2
Resource Constraints
Est. VRAM: 2GB at BF16
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Retrieval, Multilingual Retrieval, Code Retrieval
Model Details
Focus sentinel Model Family Embedding
Parameters 311M Context Window 32K License Apache 2.0 Resource Constraints Est. VRAM: 2GB at BF16 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 0.58GB 2GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-embedding-311m-multilingual-r2 at revision 44399559 on 2026-09-03. Run Granite Run Granite Embedding 311M Multilingual R2 with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-embedding-311m-multilingual-r23. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-embedding-311m-multilingual-r2", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose When retrieval quality matters more than speed; use over 97M for demanding multilingual pipelines.
Focus sentinel Granite Embedding English R2
Resource Constraints
Est. VRAM: 1GB at BF16
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Retrieval, Code Retrieval
Model Details
Focus sentinel Model Family Embedding
Parameters 149M Context Window 32K License apache-2.0 Resource Constraints Est. VRAM: 1GB at BF16 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 0.28GB 1GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-embedding-english-r2 at revision 47ea694b on 2026-09-03. Run Granite Run Granite Embedding English R2 with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-embedding-english-r23. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-embedding-english-r2", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose English-only workloads where you want a lightweight, high-quality embedding model without multilingual overhead.
Focus sentinel Granite Embedding English Small R2
Resource Constraints
Est. VRAM: 1GB at BF16
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Retrieval, Code Retrieval
Model Details
Focus sentinel Model Family Embedding
Parameters 47M Context Window 8K License apache-2.0 Resource Constraints Est. VRAM: 1GB at BF16 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 0.09GB 1GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-embedding-small-english-r2 at revision 2ab6fa8e on 2026-09-03. Run Granite Run Granite Embedding English Small R2 with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-embedding-small-english-r23. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-embedding-small-english-r2", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Ultra-lightweight English embedding model with near-200 docs/sec throughput and 8192 context — the fastest option in the Embedding family; choose the 149M English R2 if you need higher accuracy.
Focus sentinel Granite Embedding Reranker English R2
Resource Constraints
Est. VRAM: 2GB at F32
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Reranking
Model Details
Focus sentinel Model Family Embedding
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Est. VRAM: 2GB at F32 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Second-stage reranking Enterprise search Published variants Variant Format Download Est. VRAM Tier F32 (full precision)safetensors 0.56GB 2GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-embedding-reranker-english-r2 at revision d09d3d69 on 2026-09-03. Run Granite Run Granite Embedding Reranker English R2 with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-embedding-reranker-english-r23. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-embedding-reranker-english-r2", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Use as a reranker after embedding retrieval when top-result precision is critical.
Focus sentinel Resource Constraints
Est. VRAM: 18GB at BF16, down to 5GB at Q2_K
Supported GPUs: NVIDIA RTX 4090, L4, A10G
Capabilities: Harm Detection, Custom Risk/Policy Criteria (BYOC)
Model Details
Focus sentinel Model Family Guardian
Parameters 8B Context Window N/A License Apache 2.0 Resource Constraints Est. VRAM: 18GB at BF16, down to 5GB at Q2_K Supported GPUs: NVIDIA RTX 4090, L4, A10G Capabilities Harm detection Hallucination flagging Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 15.61GB 18GB Prosumer GPU (13-24GB VRAM) Q8_0 GGUF 8.3GB 10GB Consumer GPU (5-12GB VRAM) Q6_K GGUF 6.41GB 8GB Consumer GPU (5-12GB VRAM) Q5_1 GGUF 5.88GB 8GB Consumer GPU (5-12GB VRAM) Q5_K_M GGUF 5.56GB 7GB Consumer GPU (5-12GB VRAM) Q5_0 GGUF 5.42GB 7GB Consumer GPU (5-12GB VRAM) Q5_K_S GGUF 5.42GB 7GB Consumer GPU (5-12GB VRAM) Q4_1 GGUF 4.96GB 7GB Consumer GPU (5-12GB VRAM) Q4_K_M Recommended
GGUF 4.77GB 6GB Consumer GPU (5-12GB VRAM) Q4_K_S GGUF 4.53GB 6GB Consumer GPU (5-12GB VRAM) Q4_0 GGUF 4.49GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_L Heavily quantized
GGUF 4.21GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_M Heavily quantized
GGUF 3.88GB 6GB Consumer GPU (5-12GB VRAM) Q3_K_S Heavily quantized
GGUF 3.51GB 5GB Consumer GPU (5-12GB VRAM) Q2_K Heavily quantized
GGUF 3.05GB 5GB Consumer GPU (5-12GB VRAM)
Variants published as files within one repository share a link.
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-guardian-4.1-8b at revision ab01ccca on 2026-09-03. Run Granite Run Granite Guardian 4.1 8B with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-guardian-4.1-8b3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-guardian-4.1-8b", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Latest and most comprehensive; preferred for all new safety and compliance deployments.
Focus sentinel Resource Constraints
Est. VRAM: 2GB at BF16
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Document Conversion, OCR, Table Structure Recognition, Formula Recognition, Code Block Recognition, Chart-to-Table Conversion
Model Details
Focus sentinel Model Family Docling
Parameters 258M Context Window N/A License Apache 2.0 Resource Constraints Est. VRAM: 2GB at BF16 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Document conversion PDF-to-structured data Document digitization Enterprise doc workflows Published variants Variant Format Download Est. VRAM Tier BF16 (full precision)safetensors 0.48GB 2GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-docling-258M at revision 982fe3b4 on 2026-09-03. Run Granite Run Granite Docling 258M with vLLM in a container. Full guide
1. Pull the vLLM image
docker pull vllm/vllm-openai:latest2. Run the model (requires an NVIDIA GPU)
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-p 8000:8000 \
vllm/vllm-openai:latest \
--model ibm-granite/granite-docling-258M3. Send a request
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "ibm-granite/granite-docling-258M", "messages": [{"role": "user", "content": "How are you today?"}]}'Also Available On When to Choose Only model in the family; reliable document parsing with a tiny footprint.
Focus sentinel Resource Constraints
Est. VRAM: 1GB at F32
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Point and probabilistic forecasting, Multivariate Data
Model Details
Focus sentinel Model Family Time Series
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Est. VRAM: 1GB at F32 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Point and probabilistic forecasting Streaming data Multivariate data Control variables Published variants Variant Format Download Est. VRAM Tier F32 (full precision)safetensors 0.01GB 1GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-timeseries-ttm-r3 at revision ea17cfd2 on 2026-09-03. Run Granite Serving instructions for Granite TTM R3 aren't published here. See the ibm-granite/granite-timeseries-ttm-r3 model card on Hugging Face for how to run it.
When to Choose Best starting point for most forecasting tasks — extremely low resource cost and strong zero-shot baseline.
Focus sentinel Resource Constraints
Est. VRAM: 1GB at F32
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Anomaly Detection, Imputation, Similarity Search, Classification
Model Details
Focus sentinel Model Family Time Series
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Est. VRAM: 1GB at F32 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Anomaly detection Imputation Search Classification Published variants Variant Format Download Est. VRAM Tier F32 (full precision)safetensors 0GB 1GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-timeseries-tspulse-r1 at revision 2e64fcdc on 2026-09-03. Run Granite Serving instructions for Granite TSPulse R1 aren't published here. See the ibm-granite/granite-timeseries-tspulse-r1 model card on Hugging Face for how to run it.
When to Choose When you need forecasting plus anomaly detection or classification in a single compact model.
Focus sentinel Resource Constraints
Est. VRAM: 1GB at F32
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Point and probabilistic forecasting
Model Details
Focus sentinel Model Family Time Series
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Est. VRAM: 1GB at F32 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Point and probabilistic forecasting Varying input/output lengths Differing sampling rates Published variants Variant Format Download Est. VRAM Tier F32 (full precision)safetensors 0.03GB 1GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-timeseries-flowstate-r1 at revision 05effc6c on 2026-09-03. Run Granite Serving instructions for Granite FlowState R1 aren't published here. See the ibm-granite/granite-timeseries-flowstate-r1 model card on Hugging Face for how to run it.
When to Choose When TTM accuracy isn't sufficient and you can trade compute for better forecast quality.
Focus sentinel Resource Constraints
Est. VRAM: 2GB at F32
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Point and probabilistic forecasting
Model Details
Focus sentinel Model Family Time Series
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Est. VRAM: 2GB at F32 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Point and probabilistic forecasting Varying input/output lengths Published variants Variant Format Download Est. VRAM Tier F32 (full precision)safetensors 1.43GB 2GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-timeseries-patchtst-fm-r2 at revision e340dcc2 on 2026-09-03. Run Granite Serving instructions for Granite PatchTST FM R2 aren't published here. See the ibm-granite/granite-timeseries-patchtst-fm-r2 model card on Hugging Face for how to run it.
When to Choose Latest PatchTST FM generation — Conformer-based blocks, longer context (up to 8,192 timepoints), and expanded training data improve on R1's probabilistic forecasting accuracy; backward-compatible with R1 checkpoints.
Focus sentinel Resource Constraints
Est. VRAM: 2GB at F32
Runs on CPU or an entry-level GPU
Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+)
Capabilities: Point and probabilistic forecasting
Model Details
Focus sentinel Model Family Time Series
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Est. VRAM: 2GB at F32 Runs on CPU or an entry-level GPU Supported GPUs: CPU or entry-level GPU (T4, RTX 3050+) Capabilities Point and probabilistic forecasting Varying input/output lengths Published variants Variant Format Download Est. VRAM Tier F32 (full precision)safetensors 0.96GB 2GB CPU-friendly (up to 4GB VRAM)
How these estimates are calculated Inference only — fine-tuning requires roughly 3-4x this. Batch size 1, single concurrent request. KV cache sized for 4,096 tokens of context, not the model maximum. Weights measured from published file sizes; every quantization shown ships as a real published file. A model is listed if any published variant fits — the variant is named on the card. Estimates, not guarantees — actual usage varies by runtime and settings. Measured from ibm-granite/granite-timeseries-patchtst-fm-r1 at revision 151f9c6d on 2026-09-03. Run Granite Serving instructions for Granite PatchTST FM R1 aren't published here. See the ibm-granite/granite-timeseries-patchtst-fm-r1 model card on Hugging Face for how to run it.
When to Choose When you need highly accurate probabilistic forecasting for univariate time series data.
Focus sentinel Resource Constraints
Resource requirements: Not applicable
Runs on top of a base model of your choice
Capabilities: Context Attribution, Requirement Check, Uncertainty Quantification
Model Details
Focus sentinel Model Family Libraries
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Resource requirements: Not applicable Runs on top of a base model of your choice Capabilities Explainability & output verification Response confidence calibration Published variants Not applicable — GraniteLib Core R1.0 is a software library rather than a checkpoint. Its footprint is whatever model you run underneath it.
Run Granite Serving instructions for GraniteLib Core R1.0 aren't published here. See the ibm-granite/granitelib-core-r1.0 model card on Hugging Face for how to run it.
When to Choose When you want to extend a Granite base model's general capabilities without full fine-tuning.
Focus sentinel Resource Constraints
Resource requirements: Not applicable
Runs on top of a base model of your choice
Capabilities: Query Rewrite (QR), Query Clarification (QC), Context Relevance (CR), Answerability Determination (AD), Hallucination Detection (HD), Citation Generation (CG)
Model Details
Focus sentinel Model Family Libraries
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Resource requirements: Not applicable Runs on top of a base model of your choice Capabilities RAG enhancement Retrieval-tuned generation Published variants Not applicable — GraniteLib RAG R1.0 is a software library rather than a checkpoint. Its footprint is whatever model you run underneath it.
Run Granite Serving instructions for GraniteLib RAG R1.0 aren't published here. See the ibm-granite/granitelib-rag-r1.0 model card on Hugging Face for how to run it.
When to Choose Building RAG pipelines on a Granite base model and want retrieval-optimized generation behavior out of the box.
Focus sentinel Resource Constraints
Resource requirements: Not applicable
Runs on top of a base model of your choice
Capabilities: Harm Detection, Hallucination Detection (HD), Factuality Detection, Factuality Correction, Policy Guardrails
Model Details
Focus sentinel Model Family Libraries
Parameters N/A Context Window N/A License apache-2.0 Resource Constraints Resource requirements: Not applicable Runs on top of a base model of your choice Capabilities Safety layer add-on Lightweight content moderation Published variants Not applicable — GraniteLib Guardian R1.0 is a software library rather than a checkpoint. Its footprint is whatever model you run underneath it.
Run Granite Serving instructions for GraniteLib Guardian R1.0 aren't published here. See the ibm-granite/granitelib-guardian-r1.0 model card on Hugging Face for how to run it.
When to Choose When you want composable safety filtering on your existing Granite model rather than running a separate Guardian model.
Focus sentinel