flexrouter 2.3 · see what changed

Your app keeps answering when one provider runs out.

flexrouter pools the free tiers of several AI providers behind one local address.

$ pip install flexrouter

What is flexrouter?

Short answer: a tool that lets apps run on free AI without hitting the limits.

Google, Groq, Mistral and a few other AI companies give developers free access to their models, but the limits are tight. Some of Google's newest models allow 20 requests a day. Groq's free tier stops at around 30 a minute. That sounds like plenty until you realise a modern AI app, like a coding assistant or anything that uses tools, can fire off a dozen requests for a single thing you ask it. On one free account, it stalls within minutes.

flexrouter pools them. It runs in the background on your computer, and every app sends its AI requests to it instead of to one company. It tracks how much free usage is left everywhere and sends each request to the best model that still has room. When one hits its limit, the next one takes over, so the app keeps working and the bill stays at zero.

A dashboard shows where every request went, what's left for the day, and what broke and why.

Limits as of September 2026.

0/8flexrouter, open for the first time.
Overview

//Overview

Get started

0 of 5 done
  1. Add a provider
  2. Paste its key
  3. Add models with AI
  4. Test all
  5. Point your app at localhost:4891/v1

Shown on first start, and again whenever no model works. Each step ticks itself.

The dashboard

The same pages as in the story, on the real thing.

The Overview page
What the router has been doing: requests, answers, failovers and speed, hour by hour.

How it works

01

Buckets

Your app asks for a bucket like smart or fast, never a model name. Each bucket ranks its models by score or by speed, then picks at random among the ones within 20% of the best that can answer, so the load spreads out.

gemini-3.8-flash
gpt-oss-120b
mistral-medium
02

Failover

When a model is rate-limited, down or broken, the next one answers straight away, with no waiting between tries. A request pinned to one model reports the failure instead of switching, so you always know which model answered.

groqBusy
groq → googleai ↻Ready
03

Several keys per provider

Add more than one key for a provider and they are used in turn. A key that hits its limit is skipped until it resets. Keys live in their own file, never in your settings file.

groq · key 1Busy
groq · key 2Ready
04

Statuses and the error brain

Every model and key has one status: Ready, Busy, Struggling, Needs you or Off. Errors are explained in plain words. The error brain sorts each new error, says how sure it is, and asks you to check when it isn't.

ReadyBusyStrugglingNeeds youOff
05

Allowance

See what is left of every free allowance today and when each limit resets. Limits you set are marked as yours. Numbers the provider reports carry the provider's name.

gemini-3.8-flashYour limit
7,728 tokens leftGroq says
06

Setup with AI

Pick a provider from the list, marked free or paid, and paste a key. Add models with AI finds the models the key can use, Find rate limits with AI fills in their limits, and Rank models with AI puts them in order.

Free tierFree tierPaid

Built with flexrouter

stashBeing updated

An AI inventory for storage boxes. Photograph your boxes, ask what is in them, and browse them on a 3D map.

GitHub
agoraBeing updated

Two AI models argue opposite sides of a dilemma in real time while you watch.

GitHub