Llama-4-Maverick-17B-128E-Instruct-FP8

Llama-4-Maverick-17B-128E-Instruct-FP8: a llama model from llama, ~128K context, knowledge cutoff 2024-08

Published
scheduleReleasedApril 5, 2025

Llama-4-Maverick-17B-128E-Instruct-FP8 is a llama model from llama (~128K context, input around $0/1M tokens). Suitable for assistants, content generation, knowledge Q&A and business automation

starsCapabilities

visibilityVision understandingcodeFunction calling

paymentsContext and pricing

Context limit128,000
Max output4,096
Knowledge cutoff2024-08
Input price$0/ 1M tokens
Output price$0/ 1M tokens

descriptionOverview

Overview

Llama-4-Maverick-17B-128E-Instruct-FP8 is provided by llama, model ID llama-4-maverick-17b-128e-instruct-fp8

Key specs

  • Context: 128K tokens
  • Max output: 4.1K tokens
  • Knowledge cutoff: 2024-08
  • Input price: $0/1M
  • Output price: $0/1M

Best for

Consider Llama-4-Maverick-17B-128E-Instruct-FP8 when comparing context length, pricing, multimodal support and relay availability

lightbulbUse cases

  • Image and multimodal Q&A
  • Visual content analysis
  • Document and screenshot understanding

thumb_upStrengths

  • Large context window (~128K)
  • Open weights, can be self-hosted
  • Supports image and multimodal input

infoLimitations

  • Knowledge cutoff 2024-08; newer facts need external retrieval
  • Pricing and availability vary by upstream and relay; verify with official docs and tests

linkReferences

This content is compiled from official documentation and public sources. Always refer to official documentation for final details