Announcement

Usage metering and quotas for Kipchak

Mamluk · · 9 min read
Usage metering and quotas for Kipchak

Kipchak APIs can now record what each customer uses and enforce what each customer has bought. Usage Metering emits one usage event for every request, attributed to the customer who made it. Quotas enforces limits per plan, on requests and on units such as credits, after authentication. Both read the customer’s identity from Kipchak Identity, a new free library, so that a customer is the same on the invoice and in the limits.

Usage Metering and Quotas are part of Kipchak Enterprise. Kipchak Identity is free. This release also rebuilds rate limiting in the Subashi and Subashi Pro firewalls, so that limits hold under concurrent load.

Each component was tested in three stages: unit tests, integration tests against Kafka, Memcached and Valkey running in Docker, and a complete Kipchak API under FrankenPHP worker mode, with real signed tokens verified against a JWKS endpoint.

Who the customer is

Billing and quotas both need a stable name for the customer behind a request. Where that name comes from depends on how the customer authenticates. API key clients are defined in Kipchak. Clients of an identity provider are defined there, and identified by a claim in their token. Signed requests carry a key ID.

Kipchak Identity describes this once, in config/kipchak.identity.php:

return [
    'consumer' => [
        ['source' => 'token', 'claim' => 'client_id'],              // tokens from an identity provider
        ['source' => 'attribute', 'name' => 'kipchak.hmac.key_id'],  // signed requests
        ['source' => 'api_key'],                                     // API key clients
    ],
    'tenant' => [
        ['source' => 'token', 'claim' => 'org_id'],
    ],
];

Sources are tried in order, and the first that answers names the customer. An API key is never used as the name itself: the client name it maps to is used instead, so keys do not reach usage records or counters. A map on any source translates what it finds into a readable name, for example an identity provider’s opaque client IDs into customer names.

Verified tokens on the request

The JWKS and JWT authentication middlewares now place the verified token’s claims on the request, as the kipchak.auth.token attribute. Until now, they stored the token only in the dependency container.

Under FrankenPHP worker mode, the container lives for the life of the worker, not the request. In testing, 12 of 12 requests that carried no token, served after authenticated requests, found a previous caller’s token in the container. The request attribute exists only on the request that carried the verified token. The token source reads that attribute, so a customer is identified only from a token that has been verified on that request. The container entry is kept for compatibility.

Usage Metering

Usage Metering records one event for each request, in the CloudEvents 1.0 format. Each event states:

  • who made the request: the customer, and optionally the tenant;
  • what it called: the method and the route pattern, such as /v1/users/{id}, rather than the path with its identifiers;
  • how it ended: the status, the duration, and the request and response sizes;
  • what it consumed: one request, and any units the route records.

A route records units in one line:

Usage::of($request)?->add('tokens', $completion->usage->totalTokens);

Fixed costs can be configured instead, for example five credits for each report.

Events are sent to Kafka, to the application log, or to a destination of the customer’s own. Each event has a unique identifier for deduplication, and events from one customer stay in order on one Kafka partition. The customer is also the event’s CloudEvents subject, which is where usage tools that read CloudEvents look for it.

Kafka does not slow requests

A usage record must not delay the request it describes. The Kafka driver now queues messages without waiting for the broker, and delivers them from a background thread. In testing, with Kafka stopped, requests were served with a 95th-percentile response time of 5.1 milliseconds.

Events are not lost during a short outage. Sixty requests were made while Kafka was stopped, and all 60 events were delivered once it returned. If an outage lasts longer than the configured delivery timeout, each undelivered event is written to the log in full so that it can be replayed: in testing, 30 of 30 were recorded. Events still queued when a worker stops are delivered before it exits. A worker that is terminated without warning loses the events still in its queue, so the time an event waits before it is sent (the driver’s linger_ms setting) should be kept short.

Quotas

A firewall limit protects an API from abuse. A quota enforces what a customer has bought, for example 1,000 requests a day on the free plan, or 50,000 credits a month on the professional plan. The Quotas middleware runs after authentication, so it counts verified customers rather than addresses. It does not require a firewall.

Limits per plan

Plans are defined in config/kipchak.quotas.php:

return [
    'enabled' => true,
    'consumers' => ['acme-corp' => 'pro'],
    'default_plan' => 'free',
    'plans' => [
        'free' => [
            'requests' => [
                ['name' => 'burst', 'limit' => 10, 'window' => 1],
                ['name' => 'daily', 'limit' => 1000, 'period' => 'day'],
            ],
            'units' => [
                ['name' => 'monthly-credits', 'meter' => 'credits', 'limit' => 500, 'period' => 'month'],
            ],
        ],
        'pro' => [
            'requests' => [['name' => 'daily', 'limit' => 100000, 'period' => 'day']],
        ],
    ],
];
  • Windows and periods. A limit applies over a sliding window of seconds, for limits on load, or over a calendar minute, hour, day or month in UTC, for limits that are sold and invoiced.
  • Plans from the identity provider. A plan can be taken from a claim in the verified token. A customer cannot change their plan by editing a token, because only verified tokens are read.
  • Units, not only requests. Unit limits count what Usage Metering records, so the quota and the invoice use the same figures.
  • Customer, organisation or plan. A limit can count each customer, each organisation, or everyone on the plan together.
  • Clear refusals. A refused request receives status 429, a Retry-After header, the RateLimit headers defined in a current IETF draft, and a body that names the plan and the limit.

One limit for every user of an identity provider

A plan can be assigned from the token’s issuer, which identifies the identity provider. Every customer who signs in through that provider is then on the same plan, and a limit counted for the plan applies to all of them together. For example, everyone signing in through one provider can be limited to 10,000 requests an hour in total, and to 500 each.

In testing, four clients of one identity provider, sharing a limit of 20 requests, made 28 requests at the same time. Exactly 20 were accepted. Limits are checked from the narrowest to the widest, so a request refused by a customer’s own limit never takes up capacity in a shared one.

Rate limiting in Subashi and Subashi Pro

Subashi Pro is the Enterprise web application firewall for Kipchak, and Subashi is its free edition. Both run before authentication, inspecting requests and limiting abusive traffic. Their rate limiting has been rebuilt, and Quotas uses the same counters.

Limits that hold under load

Rate limits were previously counted by reading a value from the cache and writing it back. When requests arrive at the same time, that loses counts. In a test of 3,200 increments from 16 concurrent processes, the previous method recorded 450, so a limit allowed several times the configured number of requests.

Counters now use the atomic operations of Memcached and Valkey. Valkey is supported as a store. In the same test, every store recorded all 3,200 increments. In a Kipchak API with four FrankenPHP workers, 60 simultaneous requests against a limit of 20 resulted in exactly 20 accepted requests, with both Memcached and Valkey.

  • Sliding windows. A client can no longer make a full allowance of requests at the end of one window and another at the start of the next.
  • Response headers. Refused requests carry Retry-After, and responses can carry the RateLimit headers from the IETF draft.
  • Store outages. If the counter store cannot be reached, requests are allowed by default and the failure is logged, at most once a minute per worker. Previously the failure was not logged. An API can instead refuse requests while the store is unavailable.

One engine for both editions

Subashi and Subashi Pro now share one rules engine and one rate limiter, so both editions behave identically for the features they share. The improvements to rate limiting are therefore also available in the free edition. Two errors were corrected in both:

  • A query parameter sent as a list, such as ?q[]=x, caused rules that search text to fail with a server error.
  • A list rule could match a header that was absent, when the list contained an empty value.

Availability

Kipchak Identity is free:

composer require kipchak/identity

It is installed automatically with Usage Metering and Quotas. Usage Metering and Quotas are part of Kipchak Enterprise:

composer require kipchak/middleware-metering kipchak/middleware-quotas

The documentation describes each component, with examples for each authentication middleware:

All of the packages are available now.

Share
← Back to the Journal

More from the Journal