Llama 4 on Groq

Groq

· llama-4-scout-groq

Text and Language

Production

Fastest

Llama 4 served on Groq LPUs at extreme token throughput for real-time drafting and chat.

Best use cases: Real-time ideation, instant drafts, chat UX

7.5

Quality

10

Speed

9

Cost Eff.

7.5

Brand Safety

6.5

Control

7.5

Adherence

Input cost

$0.11 / 1M tokens

Output cost

$0.34 / 1M tokens

Generation cost

Max context / duration

128K tokens

Aspect ratios

Resolutions

Input types

Text

Output types

Text

API endpoint

api.groq.com/openai/v1

Commercial use

Allowed

Data retention

Zero retention

Regions available

US, EU

API connected

Connected

Last updated

Release date

DreamCache

We make machines dream.

© 2026 DreamCache AI Ltd. All rights reserved.

The Creative Innovation OS

DreamCache AI is an independent platform. Third-party model, product and company names are trademarks of their respective owners. References to providers indicate existing or planned technical compatibility and do not imply endorsement, sponsorship or formal partnership. Model and integration availability may change according to API access, commercial terms and regional restrictions.