Skip to content

FastAPI SDK ​

API key authentication, per-customer rate limiting and metered billing for your FastAPI API — powered by your MicroAuth tenant, in one dependency.

bash
pip install microauth-fastapi
python
from fastapi import FastAPI, Security
from microauth_fastapi import MicroAuth, Customer

app = FastAPI()
auth = MicroAuth(app)  # tenant secret from MICROAUTH_SECRET_KEY

@app.get("/forecast")
async def forecast(customer: Customer = Security(auth)):
    return {"hello": customer.id}

That's the whole integration. Your customers sign up on your MicroAuth developer portal, create API keys, pick a plan or top up credits — and every request to /forecast is now authenticated, rate-limited and billed.

What it does ​

  • API key auth via the X-API-Key header (configurable), including the security scheme in your OpenAPI docs — the Swagger "Authorize" button just works.
  • Suspension, balance and quota enforcement — suspended customers get 403, customers who ran out of prepaid credit get 402, customers over their monthly quota get 429.
  • Per-customer rate limiting at each customer's effective RPS (custom override → plan → pay-as-you-go default), 429 with Retry-After.
  • MicroAuth platform allowance admission using the latest monthly values from the snapshot plus local reservations. Redis coordinates reservations made by conforming SDK processes, but this is not an untrusted gateway.
  • Usage reporting and metered billing — every authenticated response is reported for request allowance and customer quota accounting. Only billable status codes (default 200 only) charge the customer balance.

Designed for the normal hot path ​

Known keys are served from the local snapshot. Network access is still needed for a key created after the latest snapshot, recovery when cached data is too old, and background synchronization:

  • A snapshot of your customers, keys and limits is cached in memory and refreshed in the background (default every 30s).
  • Keys not in the snapshot yet (created seconds ago) are resolved once via a single-flight on-demand lookup; invalid keys are negatively cached so a flood of bad keys can't reach MicroAuth.
  • Usage is put into a bounded pending queue and delivered in batches: a flush happens when 500 requests accumulate or report_interval (default 5s) elapses since the last flush, whichever comes first. Every original item keeps the same idempotency key across retries and restarts until it receives an acknowledgement.
  • Authentication itself is a SHA-256 and a couple of dict lookups.

With fail_open=False, cached data is not trusted past max_snapshot_age. The default fail_open=True can serve known keys beyond that threshold, but never beyond max_stale_snapshot_age. When the SDK cannot establish a trustworthy result, it raises a typed availability error instead of turning an outage into a false 401. fail_open never extends an expired usage_policy_id: the SDK refreshes that policy before authorizing another request and returns 503 if it cannot do so. This prevents stale pricing from producing usage that the control plane cannot safely accept.

Multiple workers? Add Redis ​

In-memory rate limiting is per-process: with 4 uvicorn workers a customer could reach roughly 4× their configured RPS. It also cannot coordinate local platform-allowance reservations with the other workers. Configure Redis when conforming processes need shared RPS or allowance state:

bash
pip install 'microauth-fastapi[redis]'
python
auth = MicroAuth(app, redis_url="redis://localhost:6379/0")

The Redis limiter uses an atomic operation and Redis server time so all workers share one counter per credential/customer and period. If Redis is unavailable, shared enforcement is unavailable. Select a conservative fallback or reject with a typed availability error according to your service policy.

Redis does not make the platform allowance a security boundary. The API owner controls the SDK process and can bypass it, and accepted traffic precedes control-plane usage reporting. Use a gateway you control if this must be a globally strict boundary against an untrusted caller.

Usage delivery is independent from rate limiting, but production reporting still needs durable pending storage and acknowledgement handling. Redis alone does not make a usage report idempotent.

Optional authentication ​

For endpoints that serve both anonymous and authenticated traffic:

python
@app.get("/status")
async def status(customer: Customer | None = Security(auth.optional)):
    return {"authenticated": customer is not None}

The Customer object ​

The dependency resolves to a Customer with the fields your handler might need:

FieldMeaning
idThe customer's MicroAuth ID (stable UUID)
key_idThe API key that authenticated this request
statusactive (suspended customers are rejected before your handler)
billing_modelThe customer's effective billing model (payg, plan, none, …)
rpsEffective requests-per-second limit
price_per_request_microEffective per-request price in micro-USD
monthly_quotaEffective monthly request cap, or None for uncapped
credit_balance_microLocally estimated prepaid balance in micro-USD after this request's reservation

Management and portal billing responses use plan_id: str | None and plan_name: str | None. plan_id is a UUID string, never an integer. The SDK principal exposes resolved limits so handlers should branch on billing_model or other effective fields rather than parsing a plan name.

Settings ​

Everything has a sensible default; override only what you need.

SettingDefaultWhat it does
secret_key$MICROAUTH_SECRET_KEYTenant secret key (mas_...)
base_urlhttps://api.microauth.comMicroAuth API ($MICROAUTH_BASE_URL)
header_nameX-API-KeyHeader customers send their key in
redis_url$MICROAUTH_REDIS_URLCoordinates RPS and request reservations across processes
sync_interval30Seconds between snapshot refreshes
report_interval5Flush when 500 requests accumulate or this many seconds pass since the last flush
flush_on_responseAuto on Vercel/LambdaEvaluate the batching rule after the final response frame
trailing_flushFalseHold a serverless invocation until the batching deadline so the last burst before traffic stops is delivered
max_snapshot_age300Staleness threshold that triggers 503 when fail_open=False
max_stale_snapshot_age3 × max_snapshot_ageAbsolute ceiling for serving known keys from stale data
fail_openTrueKeep serving known cached keys until the absolute stale ceiling
enforce_balanceTrue402 when prepaid credit is exhausted
enforce_quotaTrue429 when the monthly quota is used up
enforce_rpsTruePer-customer RPS limiting
enforce_platform_allowanceTrueReserve against the latest platform allowance before the handler
verify_negative_ttl30Seconds an invalid key is cached
timeout5HTTP timeout for MicroAuth calls
usage_spool_pathTenant-scoped temporary pathAppend-only journal that keeps in-flight usage and stable retry IDs across restarts
persistence_namespaceNoneExplicit stable Redis and default-spool scope when tenant secrets rotate
persist_usageTruePreserve unacknowledged usage across restarts
max_usage_queue10000Maximum reserved and queued usage items
shutdown_timeout10Seconds allowed for the final usage drain
redis_clientNoneExternally owned Redis client, see "Redis connections under load"

Tuning for serverless deployments ​

Serverless runtimes break two assumptions the defaults are built on: the process does not stay alive between requests, and there is no graceful shutdown. Vercel and AWS Lambda freeze or discard an instance as soon as its work is done, so background timers stop firing and files written to the instance's disk can vanish with it. The SDK detects both platforms through their environment variables and adapts where it safely can, but a few settings deserve an explicit decision.

A solid starting point for Vercel with Fluid Compute, with MICROAUTH_SECRET_KEY and MICROAUTH_REDIS_URL set in the project's environment:

python
auth = MicroAuth(app, trailing_flush=True)

Everything else can stay at its default. Here is why Redis and the trailing flush matter, and what the surrounding settings do on a serverless platform.

redis_url. Treat it as required. Without Redis, completed usage is journaled to the instance's own filesystem, and an instance that gets replaced takes its unreported usage with it. With Redis, usage enters a durable shared queue before the response is released, and any instance can recover and deliver another instance's events. Snapshots are shared too, so a cold start reuses a validated snapshot instead of hitting the control plane, and rate limits are enforced across all instances instead of per instance. Managed offerings such as Upstash persist by default. One practical note: confirm the variable is actually set in the production environment of your hosting platform. A misspelled name does not raise an error; the SDK silently falls back to per-instance behavior. Serverless platforms also multiply Redis connections in ways that can exhaust a managed plan's connection cap during bursts; the next section covers how to bound the pool.

flush_on_response. Resolves to True automatically when the VERCEL or AWS_LAMBDA_FUNCTION_NAME environment variable is present, so there is usually nothing to configure. After the final response frame is sent, the SDK checks whether a batch is due (500 requests, or report_interval seconds since the last flush) and delivers it while the invocation is still active. Without this, delivery would depend entirely on a background timer that a frozen instance never runs.

trailing_flush. Off by default; turn it on for Vercel with Fluid Compute. The response-bound check above only ships batches that are already due. A burst that finishes inside the interval leaves its usage queued, and with no further traffic nothing ever delivers it, which shows up as billing that lags until the next visitor arrives. With trailing_flush=True, the last response of a burst keeps the invocation alive until the batching deadline (at most report_interval seconds) and then delivers. Concurrent responses share one waiter and it makes a single delivery attempt, so MicroAuth traffic does not increase; the only cost is a few extra seconds of instance time after a burst. Callers are unaffected because the response has already been sent. Leave it off on platforms that buffer the entire response before returning it (for example Lambda behind an adapter without response streaming), where the hold would land on your callers instead.

report_interval. On serverless this bounds two things at once: how far billing can lag behind traffic, and how long a trailing hold can last. The default of 5 seconds is a sensible middle. Lowering it reports sooner and shortens holds at the price of more frequent usage calls; raising it does the reverse.

timeout. Keep the default of 5 seconds. Every control-plane call makes up to three attempts, so the worst case is roughly three times this value, and the first request after a cold start fetches a snapshot on the caller's critical path. A generous timeout such as 30 seconds turns an unhealthy control plane into a first request that stalls for a minute and a half and collides with your platform's function duration limit.

shutdown_timeout and aclose(). Do not plan around them here. Serverless platforms rarely deliver a graceful shutdown, so the final drain may simply never run. Durability has to come from the Redis queue and from delivering usage while invocations are still active, which is one more reason to configure Redis and enable the trailing flush.

Redis connections under load ​

When you pass only redis_url, the SDK builds its own client with one-second socket timeouts and the redis-py default connection pool. That pool has no practical ceiling: any command that cannot find a free connection opens a new one. On a long-lived server with steady traffic this settles into a handful of connections and never bothers anyone. A traffic burst behaves differently. One process handling a hundred concurrent requests can briefly hold dozens of connections, and on serverless platforms every instance brings its own pool, so the totals multiply exactly when your Redis provider's connection cap starts to matter. Managed plans commonly cap connections at a few dozen on smaller tiers, and every fresh connection pays a TLS handshake that has to fit inside the one-second connect timeout. Hit either limit and commands start failing.

The SDK degrades deliberately when that happens, but not for free. Watch for these two warnings:

microauth: durable usage queue sweep failed (...)
microauth: durable usage queue is unavailable (...); falling back to the local journal

The first is a failed recovery pass; it retries on the next interval and loses nothing by itself. The second one deserves attention on serverless: usage is falling back to the instance's local journal, and if that instance is recycled before the events deliver, they are gone. Billing that quietly runs behind traffic usually starts with one of these lines.

The remedy is to own the client and bound it. A BlockingConnectionPool makes requests wait briefly for a free connection instead of opening new ones or failing outright:

python
import os
from redis.asyncio import Redis, BlockingConnectionPool

pool = BlockingConnectionPool.from_url(
    os.environ["MICROAUTH_REDIS_URL"],
    max_connections=10,
    timeout=5,
    socket_timeout=2.0,
    socket_connect_timeout=2.0,
)
auth = MicroAuth(app, trailing_flush=True, redis_client=Redis(connection_pool=pool))

Treat those numbers as a starting point, not a prescription. The right values depend on your traffic shape:

  • max_connections is a per-process budget. Multiply it by the number of instances or workers you realistically run at peak and keep the product comfortably under your Redis plan's connection cap, leaving headroom for dashboards and other clients. Ten per instance suits bursty serverless traffic; a busy long-lived server with high steady concurrency may deserve more.
  • timeout is how long a request waits for a free connection before giving up. A few seconds is plenty. If waiting starts showing up in your latency percentiles, the pool is too small for your concurrency: raise the budget or the plan cap rather than the timeout.
  • The socket timeouts trade tolerance for tail latency. Two seconds absorbs TLS handshakes and the occasional slow round trip. If you need more than that, look at the real problem first: keep Redis in the same region as your API.

Whatever you choose, verify it against your own load rather than trusting defaults, ours or these. Replay your peak burst against a staging deployment and watch two places: your provider's connection and error metrics, and your logs for the two warnings above. If neither complains at your real traffic shape, the pool fits.

One lifecycle note: a client you pass in is externally owned. The SDK uses it but never closes it, so close it in your own shutdown path if your platform gives you one.

Error responses ​

StatusWhen
401Missing or invalid API key
402Prepaid balance exhausted
403Customer suspended
429RPS, customer monthly quota or MicroAuth platform allowance exceeded
503Snapshot, key verification or required shared state is unavailable

All of these subclass microauth_fastapi.AuthDenied (itself a FastAPI HTTPException), so you can add your own exception handler to reshape the response bodies:

python
from microauth_fastapi import AuthDenied

@app.exception_handler(AuthDenied)
async def auth_denied(request, exc):
    return JSONResponse(
        status_code=exc.status_code,
        content={"error": exc.detail, "docs": "https://docs.yourapi.com/errors"},
        headers=exc.headers or {},
    )

Delivery and consistency notes ​

  • Balance and customer quota enforcement use the cached snapshot plus local activity. The platform allowance uses the cached snapshot plus a pre-handler reservation. These are honest-client admission controls, not authoritative gateway guarantees.
  • The reporter removes an item after accepted or duplicate. An actionable rejected result is also terminal and releases its local reservation; retry, transport failures, and missing item results stay pending.
  • Every request captures the immutable usage_policy_id returned with its customer snapshot. Delayed reports therefore use the pricing, quota, billable statuses and platform allowance that authorized that request, even if the tenant changes those settings before delivery.
  • Await auth.aclose() during graceful shutdown. If the durable queue cannot drain, shutdown surfaces the remaining work instead of claiming it was delivered.
  • Suspensions and key revocations propagate within one sync_interval (30s by default) to every worker.

Not using FastAPI? ​

The SDK is a thin, well-behaved client of three HTTP endpoints. The custom integration guide documents them so you can implement the same pattern in Go, Node, Rails or anything else.

MicroAuth is a product of Zyref, LLC.