Llama 4 on Groq
Groq
· llama-4-scout-groq
Text and Language
Production
Fastest
Llama 4 served on Groq LPUs at extreme token throughput for real-time drafting and chat.
Best use cases: Real-time ideation, instant drafts, chat UX
7.5
Quality
10
Speed
9
Cost Eff.
7.5
Brand Safety
6.5
Control
7.5
Adherence
Input cost
$0.11 / 1M tokens
Output cost
$0.34 / 1M tokens
Generation cost
—
Max context / duration
128K tokens
Aspect ratios
—
Resolutions
—
Input types
Text
Output types
Text
API endpoint
api.groq.com/openai/v1
Commercial use
Allowed
Data retention
Zero retention
Regions available
US, EU
API connected
Connected
Last updated
Release date
Documentation