Connect Meta Llama

Via a host

No direct connector, because there is nothing to connect: Meta does not sell inference. Llama is open-weight, so your Llama spend arrives on the bill of whichever host serves it — and Ratelytics ingests those hosts.

Steps

  1. Work out where your Llama tokens are actually served. The invoice you pay is the host's, not Meta's.
  2. OpenRouter (live connector): one key covers Llama alongside every other model you route through it — the right choice if you already aggregate there.
  3. Fireworks (live connector): first-party Llama serving with per-model usage; pick this if Fireworks is your primary host.
  4. Groq (CSV preset): fill the Ratelytics template from your Groq console figures — Groq has no usage API to pull.
  5. Together AI (CSV preset): same pattern, using the Together template.
  6. Self-hosted Llama (your own GPUs) has no per-token bill to ingest; track the fixed infrastructure cost in the Subscriptions registry instead.

Good to know

  • Pick ONE route per workload. The same Llama tokens showing up through both a gateway and the host underneath would count the money twice.
  • Spend appears in Ratelytics under the host's name, with the Llama model on each row — filter by model to see Llama across every host at once.

Official documentation: www.llama.com