FastAPI SDK
API key authentication, per-customer rate limiting and metered billing for your FastAPI API — powered by your MicroAuth tenant, in one dependency.
pip install microauth-fastapifrom fastapi import FastAPI, Security
from microauth_fastapi import MicroAuth, Customer
app = FastAPI()
auth = MicroAuth(app) # tenant secret from MICROAUTH_SECRET_KEY
@app.get("/forecast")
async def forecast(customer: Customer = Security(auth)):
return {"hello": customer.id}That's the whole integration. Your customers sign up on your MicroAuth developer portal, create API keys, pick a plan or top up credits — and every request to /forecast is now authenticated, rate-limited and billed.
What it does
- API key auth via the
X-API-Keyheader (configurable), including the security scheme in your OpenAPI docs — the Swagger "Authorize" button just works. - Suspension, balance and quota enforcement — suspended customers get
403, customers who ran out of prepaid credit get402, customers over their monthly quota get429. - Per-customer rate limiting at each customer's effective RPS (custom override → plan → pay-as-you-go default),
429withRetry-After. - MicroAuth platform allowance admission using the latest monthly values from the snapshot plus local reservations. Redis coordinates reservations made by conforming SDK processes, but this is not an untrusted gateway.
- Usage reporting and metered billing — every authenticated response is reported for request allowance and customer quota accounting. Only billable status codes (default
200only) charge the customer balance.
Designed for the normal hot path
Known keys are served from the local snapshot. Network access is still needed for a key created after the latest snapshot, recovery when cached data is too old, and background synchronization:
- A snapshot of your customers, keys and limits is cached in memory and refreshed in the background (default every 30s).
- Keys not in the snapshot yet (created seconds ago) are resolved once via a single-flight on-demand lookup; invalid keys are negatively cached so a flood of bad keys can't reach MicroAuth.
- Usage is put into a bounded pending queue and delivered in batches: a flush happens when 500 requests accumulate or
report_interval(default 5s) elapses since the last flush, whichever comes first. Every original item keeps the same idempotency key across retries and restarts until it receives an acknowledgement. - Authentication itself is a SHA-256 and a couple of dict lookups.
With fail_open=False, cached data is not trusted past max_snapshot_age. The default fail_open=True can serve known keys beyond that threshold, but never beyond max_stale_snapshot_age. When the SDK cannot establish a trustworthy result, it raises a typed availability error instead of turning an outage into a false 401. fail_open never extends an expired usage_policy_id: the SDK refreshes that policy before authorizing another request and returns 503 if it cannot do so. This prevents stale pricing from producing usage that the control plane cannot safely accept.
Multiple workers? Add Redis
In-memory rate limiting is per-process: with 4 uvicorn workers a customer could reach roughly 4× their configured RPS. It also cannot coordinate local platform-allowance reservations with the other workers. Configure Redis when conforming processes need shared RPS or allowance state:
pip install 'microauth-fastapi[redis]'auth = MicroAuth(app, redis_url="redis://localhost:6379/0")The Redis limiter uses an atomic operation and Redis server time so all workers share one counter per credential/customer and period. If Redis is unavailable, shared enforcement is unavailable. Select a conservative fallback or reject with a typed availability error according to your service policy.
Redis does not make the platform allowance a security boundary. The API owner controls the SDK process and can bypass it, and accepted traffic precedes control-plane usage reporting. Use a gateway you control if this must be a globally strict boundary against an untrusted caller.
Usage delivery is independent from rate limiting, but production reporting still needs durable pending storage and acknowledgement handling. Redis alone does not make a usage report idempotent.
Optional authentication
For endpoints that serve both anonymous and authenticated traffic:
@app.get("/status")
async def status(customer: Customer | None = Security(auth.optional)):
return {"authenticated": customer is not None}The Customer object
The dependency resolves to a Customer with the fields your handler might need:
| Field | Meaning |
|---|---|
id | The customer's MicroAuth ID (stable UUID) |
key_id | The API key that authenticated this request |
status | active (suspended customers are rejected before your handler) |
billing_model | The customer's effective billing model (payg, plan, none, …) |
rps | Effective requests-per-second limit |
price_per_request_micro | Effective per-request price in micro-USD |
monthly_quota | Effective monthly request cap, or None for uncapped |
credit_balance_micro | Locally estimated prepaid balance in micro-USD after this request's reservation |
Management and portal billing responses use plan_id: str | None and plan_name: str | None. plan_id is a UUID string, never an integer. The SDK principal exposes resolved limits so handlers should branch on billing_model or other effective fields rather than parsing a plan name.
Settings
Everything has a sensible default; override only what you need.
| Setting | Default | What it does |
|---|---|---|
secret_key | $MICROAUTH_SECRET_KEY | Tenant secret key (mas_...) |
base_url | https://api.microauth.com | MicroAuth API ($MICROAUTH_BASE_URL) |
header_name | X-API-Key | Header customers send their key in |
redis_url | $MICROAUTH_REDIS_URL | Coordinates RPS and request reservations across processes |
sync_interval | 30 | Seconds between snapshot refreshes |
report_interval | 5 | Flush when 500 requests accumulate or this many seconds pass since the last flush |
flush_on_response | Auto on Vercel/Lambda | Evaluate the batching rule after the final response frame |
trailing_flush | False | Hold a serverless invocation until the batching deadline so the last burst before traffic stops is delivered |
max_snapshot_age | 300 | Staleness threshold that triggers 503 when fail_open=False |
max_stale_snapshot_age | 3 × max_snapshot_age | Absolute ceiling for serving known keys from stale data |
fail_open | True | Keep serving known cached keys until the absolute stale ceiling |
enforce_balance | True | 402 when prepaid credit is exhausted |
enforce_quota | True | 429 when the monthly quota is used up |
enforce_rps | True | Per-customer RPS limiting |
enforce_platform_allowance | True | Reserve against the latest platform allowance before the handler |
verify_negative_ttl | 30 | Seconds an invalid key is cached |
timeout | 5 | HTTP timeout for MicroAuth calls |
usage_spool_path | Tenant-scoped temporary path | Append-only journal that keeps in-flight usage and stable retry IDs across restarts |
persistence_namespace | None | Explicit stable Redis and default-spool scope when tenant secrets rotate |
persist_usage | True | Preserve unacknowledged usage across restarts |
max_usage_queue | 10000 | Maximum reserved and queued usage items |
shutdown_timeout | 10 | Seconds allowed for the final usage drain |
redis_client | None | Externally owned Redis client, see "Redis connections under load" |
Tuning for serverless deployments
Serverless runtimes break two assumptions the defaults are built on: the process does not stay alive between requests, and there is no graceful shutdown. Vercel and AWS Lambda freeze or discard an instance as soon as its work is done, so background timers stop firing and files written to the instance's disk can vanish with it. The SDK detects both platforms through their environment variables and adapts where it safely can, but a few settings deserve an explicit decision.
A solid starting point for Vercel with Fluid Compute, with MICROAUTH_SECRET_KEY and MICROAUTH_REDIS_URL set in the project's environment:
auth = MicroAuth(app, trailing_flush=True)Everything else can stay at its default. Here is why Redis and the trailing flush matter, and what the surrounding settings do on a serverless platform.
redis_url. Treat it as required. Without Redis, completed usage is journaled to the instance's own filesystem, and an instance that gets replaced takes its unreported usage with it. With Redis, usage enters a durable shared queue before the response is released, and any instance can recover and deliver another instance's events. Snapshots are shared too, so a cold start reuses a validated snapshot instead of hitting the control plane, and rate limits are enforced across all instances instead of per instance. Managed offerings such as Upstash persist by default. One practical note: confirm the variable is actually set in the production environment of your hosting platform. A misspelled name does not raise an error; the SDK silently falls back to per-instance behavior. Serverless platforms also multiply Redis connections in ways that can exhaust a managed plan's connection cap during bursts; the next section covers how to bound the pool.
flush_on_response. Resolves to True automatically when the VERCEL or AWS_LAMBDA_FUNCTION_NAME environment variable is present, so there is usually nothing to configure. After the final response frame is sent, the SDK checks whether a batch is due (500 requests, or report_interval seconds since the last flush) and delivers it while the invocation is still active. Without this, delivery would depend entirely on a background timer that a frozen instance never runs.
trailing_flush. Off by default; turn it on for Vercel with Fluid Compute. The response-bound check above only ships batches that are already due. A burst that finishes inside the interval leaves its usage queued, and with no further traffic nothing ever delivers it, which shows up as billing that lags until the next visitor arrives. With trailing_flush=True, the last response of a burst keeps the invocation alive until the batching deadline (at most report_interval seconds) and then delivers. Concurrent responses share one waiter and it makes a single delivery attempt, so MicroAuth traffic does not increase; the only cost is a few extra seconds of instance time after a burst. Callers are unaffected because the response has already been sent. Leave it off on platforms that buffer the entire response before returning it (for example Lambda behind an adapter without response streaming), where the hold would land on your callers instead.
report_interval. On serverless this bounds two things at once: how far billing can lag behind traffic, and how long a trailing hold can last. The default of 5 seconds is a sensible middle. Lowering it reports sooner and shortens holds at the price of more frequent usage calls; raising it does the reverse.
timeout. Keep the default of 5 seconds. Every control-plane call makes up to three attempts, so the worst case is roughly three times this value, and the first request after a cold start fetches a snapshot on the caller's critical path. A generous timeout such as 30 seconds turns an unhealthy control plane into a first request that stalls for a minute and a half and collides with your platform's function duration limit.
shutdown_timeout and aclose(). Do not plan around them here. Serverless platforms rarely deliver a graceful shutdown, so the final drain may simply never run. Durability has to come from the Redis queue and from delivering usage while invocations are still active, which is one more reason to configure Redis and enable the trailing flush.
Redis connections under load
When you pass only redis_url, the SDK builds its own client with one-second socket timeouts and the redis-py default connection pool. That pool has no practical ceiling: any command that cannot find a free connection opens a new one. On a long-lived server with steady traffic this settles into a handful of connections and never bothers anyone. A traffic burst behaves differently. One process handling a hundred concurrent requests can briefly hold dozens of connections, and on serverless platforms every instance brings its own pool, so the totals multiply exactly when your Redis provider's connection cap starts to matter. Managed plans commonly cap connections at a few dozen on smaller tiers, and every fresh connection pays a TLS handshake that has to fit inside the one-second connect timeout. Hit either limit and commands start failing.
The SDK degrades deliberately when that happens, but not for free. Watch for these two warnings:
microauth: durable usage queue sweep failed (...)
microauth: durable usage queue is unavailable (...); falling back to the local journalThe first is a failed recovery pass; it retries on the next interval and loses nothing by itself. The second one deserves attention on serverless: usage is falling back to the instance's local journal, and if that instance is recycled before the events deliver, they are gone. Billing that quietly runs behind traffic usually starts with one of these lines.
The remedy is to own the client and bound it. A BlockingConnectionPool makes requests wait briefly for a free connection instead of opening new ones or failing outright:
import os
from redis.asyncio import Redis, BlockingConnectionPool
pool = BlockingConnectionPool.from_url(
os.environ["MICROAUTH_REDIS_URL"],
max_connections=10,
timeout=5,
socket_timeout=2.0,
socket_connect_timeout=2.0,
)
auth = MicroAuth(app, trailing_flush=True, redis_client=Redis(connection_pool=pool))Treat those numbers as a starting point, not a prescription. The right values depend on your traffic shape:
max_connectionsis a per-process budget. Multiply it by the number of instances or workers you realistically run at peak and keep the product comfortably under your Redis plan's connection cap, leaving headroom for dashboards and other clients. Ten per instance suits bursty serverless traffic; a busy long-lived server with high steady concurrency may deserve more.timeoutis how long a request waits for a free connection before giving up. A few seconds is plenty. If waiting starts showing up in your latency percentiles, the pool is too small for your concurrency: raise the budget or the plan cap rather than the timeout.- The socket timeouts trade tolerance for tail latency. Two seconds absorbs TLS handshakes and the occasional slow round trip. If you need more than that, look at the real problem first: keep Redis in the same region as your API.
Whatever you choose, verify it against your own load rather than trusting defaults, ours or these. Replay your peak burst against a staging deployment and watch two places: your provider's connection and error metrics, and your logs for the two warnings above. If neither complains at your real traffic shape, the pool fits.
One lifecycle note: a client you pass in is externally owned. The SDK uses it but never closes it, so close it in your own shutdown path if your platform gives you one.
Error responses
| Status | When |
|---|---|
401 | Missing or invalid API key |
402 | Prepaid balance exhausted |
403 | Customer suspended |
429 | RPS, customer monthly quota or MicroAuth platform allowance exceeded |
503 | Snapshot, key verification or required shared state is unavailable |
All of these subclass microauth_fastapi.AuthDenied (itself a FastAPI HTTPException), so you can add your own exception handler to reshape the response bodies:
from microauth_fastapi import AuthDenied
@app.exception_handler(AuthDenied)
async def auth_denied(request, exc):
return JSONResponse(
status_code=exc.status_code,
content={"error": exc.detail, "docs": "https://docs.yourapi.com/errors"},
headers=exc.headers or {},
)Delivery and consistency notes
- Balance and customer quota enforcement use the cached snapshot plus local activity. The platform allowance uses the cached snapshot plus a pre-handler reservation. These are honest-client admission controls, not authoritative gateway guarantees.
- The reporter removes an item after
acceptedorduplicate. An actionablerejectedresult is also terminal and releases its local reservation;retry, transport failures, and missing item results stay pending. - Every request captures the immutable
usage_policy_idreturned with its customer snapshot. Delayed reports therefore use the pricing, quota, billable statuses and platform allowance that authorized that request, even if the tenant changes those settings before delivery. - Await
auth.aclose()during graceful shutdown. If the durable queue cannot drain, shutdown surfaces the remaining work instead of claiming it was delivered. - Suspensions and key revocations propagate within one
sync_interval(30s by default) to every worker.
Not using FastAPI?
The SDK is a thin, well-behaved client of three HTTP endpoints. The custom integration guide documents them so you can implement the same pattern in Go, Node, Rails or anything else.