Skip to content

Incident detail

Model Performance Issue: llama-3.1-8b-instant

Resolved incidentMinor

Timeline window

to

Outage alerts

Get alerted the next time Groq breaks

Free email alerts for up to 5 providers. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update Groq posted, oldest to newest, exactly as it appeared on their official status page.

  1. Identified

    We have identified an issue affecting the lama-3.1-8b-instant service. A fix is in progress. Users may still experience latency until the fix completes.

  2. Monitoring

    We have implemented a fix for llama-3.1-8b-instant service and performance is improving. We’re now monitoring the model to ensure stability. If all remains normal, we will resolve the incident in the next update.

  3. Resolved

    The issues affecting llama-3.1-8b-instant have been resolved. The model is operating normally. Root cause: Two recent changes to the inference-engine-instances repository that were contributing to the elevated Orion LoRA latencies were reverted . We apologize for the disruption and thank you for your patience.

Keep exploring

More from Groq

Neighboring incidents on Groq's timeline and the rest of their record on OutageDeck.