Llama-4-Scout-17B-16E-Instruct-FP8
Llama-4-Scout-17B-16E-Instruct-FP8: a llama model from llama, ~128K context, knowledge cutoff 2024-08
Llama-4-Scout-17B-16E-Instruct-FP8 is a llama model from llama (~128K context, input around $0/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation
starsCapabilities
visibilityVision understandingcodeFunction calling
paymentsContext and pricing
Context limit128,000
Max output4,096
Knowledge cutoff2024-08
Input price$0/ 1M tokens
Output price$0/ 1M tokens
descriptionOverview
Overview
Llama-4-Scout-17B-16E-Instruct-FP8 is provided by llama, model ID llama-4-scout-17b-16e-instruct-fp8
Key specs
- Context: 128K tokens
- Max output: 4.1K tokens
- Knowledge cutoff: 2024-08
- Input price: $0/1M
- Output price: $0/1M
Best for
Consider Llama-4-Scout-17B-16E-Instruct-FP8 when comparing context length, pricing, multimodal support and relay availability
lightbulbUse cases
- Image and multimodal Q&A
- Visual content analysis
- Document and screenshot understanding
thumb_upStrengths
- Large context window (~128K)
- Open weights, can be self-hosted
- Supports image and multimodal input
infoLimitations
- Knowledge cutoff 2024-08; newer facts need external retrieval
- Pricing and availability vary by upstream and relay; verify with official docs and tests
Scan to join WeChat group