Metadata-Version: 2.3
Name: abilian-cdn
Version: 0.1.0
Summary: A small CDN: S3-backed storage and pull zones, tokens, a file browser, and download stats.
Author: Stefane Fermigier
Author-email: Stefane Fermigier <sf@abilian.com>
Requires-Dist: advanced-alchemy[litestar]>=1.11.0
Requires-Dist: aiobotocore>=3.9.1
Requires-Dist: asyncpg>=0.31.0
Requires-Dist: cyclopts>=4.25.2
Requires-Dist: httpx>=0.28.1
Requires-Dist: litestar[standard]>=2.24.0
Requires-Dist: pydantic>=2.13.5
Requires-Dist: sqlalchemy>=2.0.52
Requires-Dist: uvicorn>=0.52.4
Requires-Python: >=3.12
Description-Content-Type: text/markdown

# Abilian CDN

**Put your files somewhere fast, get a URL, see who downloads them.**

This is a small content delivery service you run yourself. Files live in S3 (or MinIO, Garage, or Scaleway), get cached on the server's local disk, and are served over HTTP with proper caching headers, byte ranges, signed URLs and download statistics. One server, not a distributed network: a request lands on one machine and is answered there.

It exists because "we need somewhere to put the release tarballs / the product images / the client's video, with a stable URL and no surprises on the bill" comes up constantly. The usual answers are either a bucket with an ugly URL and no statistics, or a CDN account with a contract attached.

**What you get:**

- **A stable public URL per file**, served with `ETag`, `Last-Modified`, `Cache-Control` and byte ranges, so browsers, `curl -C -` and download managers all behave.
- **Upload from anywhere**: a CLI for CI, drag-and-drop in the browser, or plain HTTP `PUT`. Large files are fine: the cap is 2 GiB per upload and is set by a config key.
- **Private files behind time-limited links** you can hand out, without making the whole zone private or minting a new credential.
- **Download statistics per file**, per day, with a cache hit ratio and a CSV export.
- **Mirror a site you already run** (a "pull zone"), and keep serving it when that site goes down.
- **Tools you already have.** The service exposes an S3-compatible endpoint, so `rclone`, the AWS CLI and anything built on boto can push to a zone without a plugin or a client of ours.
- **Several tenants on one server.** Organisations, users and tokens, each seeing only their own.

---

## Try it in five minutes

You need PostgreSQL, [uv](https://docs.astral.sh/uv/), and the [Garage](https://garagehq.deuxfleurs.fr/) binary on your `PATH` (`brew install garage`) standing in for S3.

```bash
./scripts/garage-dev.sh start        # a single-node S3 under var/garage/
createdb abilian_cdn_dev

source var/garage/env                # the S3_* credentials that script wrote
export DATABASE_URL=postgresql:///abilian_cdn_dev
export CDN_BASE_URL=http://localhost:8000
export CDN_SECRET_KEY=dev-secret-change-me     # any stable string will do here
export CDN_CACHE_DIR=var/cache

uv run cdn migrate
uv run cdn create org acme "ACME Corp"
uv run cdn create zone assets --org acme
uv run cdn create user you@example.com "Your Name" --org acme --role owner
uv run cdn create token laptop --org acme --scopes read,write,delete,purge,stats
```

The last two commands print a password and a token, once each. Keep them.

`CDN_SECRET_KEY` signs session and CSRF cookies, so changing it signs everyone out. Keep it stable in development and secret in production (`openssl rand -hex 32`).

```bash
uv run uvicorn abilian_cdn.app:serve --factory --port 8000
```

In another terminal:

```bash
export CDN_URL=http://localhost:8000
export CDN_TOKEN=cdn_…             # the token printed above

echo "hello from the CDN" > hello.txt
uv run cdn put assets hello.txt    # prints http://localhost:8000/assets/hello.txt

curl -i http://localhost:8000/assets/hello.txt    # look at X-Cache, ETag, Cache-Control
curl -i http://localhost:8000/assets/hello.txt    # X-Cache: HIT
```

Then open <http://localhost:8000/> and sign in with the email and password from `create user`: your zones, a file browser you can drop files onto, per-file download counts, and the traffic charts.

> **Don't point a dev server at `abilian_cdn_test`**: the test suite empties that database when it starts.

---

## Publishing files

```bash
cdn login https://cdn.example.com cdn_xxxxxxxxxxxx_…   # or set CDN_URL / CDN_TOKEN
cdn put assets ./logo.png images/logo.png              # prints the public URL
cdn sync assets ./site [--delete] [--dry-run]          # uploads only what differs
cdn ls assets[/prefix] [--long]
cdn rm assets images/logo.png
cdn url assets images/logo.png [--expires 2h]          # a plain or a signed URL
cdn purge assets images/logo.png                       # or --prefix images/, or --all
```

`sync` is the one CI will call. It compares each file's MD5 against what the zone already holds, so re-running it after a fresh checkout uploads nothing; comparing modification times would re-upload everything. The credential goes to `$XDG_CONFIG_HOME/abilian-cdn/config.toml` with mode 0600; `CDN_URL` and `CDN_TOKEN` override it, which is what a CI job should use.

A **purge** forgets a cached copy; it never destroys anything. The next request fetches the file again (from S3 for a storage zone, from the origin for a pull zone). There is a purge form on each zone's settings page too; emptying a whole zone asks you to type its name first.

---

## Using it from an S3 client

A zone is a bucket and an object path is a key, so anything that speaks S3 can talk to it:

```bash
cdn create token ci --org acme --scopes read,write,delete --s3
# prints an access key id, a secret, and the endpoint to configure

aws --endpoint-url https://cdn.example.com/_/s3 s3 sync ./dist s3://assets/
rclone sync ./dist cdn:assets          # with the endpoint set in rclone.conf
```

**A hostname of its own, for clients that need one.** Set `CDN_S3_HOST=s3.cdn.example.com` and that hostname serves the S3 API at its root (`https://s3.cdn.example.com/assets/logo.png`) while every other hostname keeps serving the CDN. It is optional. Only one client needs it: `s3cmd` folds an endpoint's path into the `Host` header it signs, so it can never authenticate against `/_/s3`. With the hostname set, s3cmd signs the right `Host` header and the request succeeds:

```bash
s3cmd --host=s3.cdn.example.com --host-bucket= put ./logo.png s3://assets/logo.png
```

(An empty `--host-bucket=` puts the bucket in the path.)

`GetObject`, `HeadObject`, `PutObject`, `CopyObject`, `DeleteObject`, `DeleteObjects`, `ListObjects` (both versions), `ListBuckets`, `HeadBucket` and `CreateBucket` are implemented, with `x-amz-meta-*` stored and returned; anything else answers `501 NotImplemented` in the XML a client can explain. `CreateBucket` never creates anything. Buckets here are zones; an administrator makes those. A zone that is already yours answers 200, which is what S3 answers for a bucket you already own and what `rclone` checks before every sync. A key can do only what its token's scopes allow, so a read-only token stays read-only however it is used. Requests are authenticated with AWS Signature Version 4.

**The S3 secret is not the bearer token's secret.** SigV4 is checked by recomputing the client's HMAC, so the service has to know the secret. Each S3 credential is derived from the service key and the token's key id, so the database holds nothing reversible to steal; rotating `CDN_SECRET_KEY` invalidates every S3 credential exactly as it invalidates every session. `cdn create token --s3` prints it.

`rclone sync` is verified end to end: uploads a tree, is idempotent on a second run, deletes what is gone, and `rclone check` reports no differences.

**Three limits.** There is no multipart upload. A single `PUT` is good for anything up to `CDN_MAX_UPLOAD_BYTES` (2 GiB by default), but clients switch to multipart above a threshold of their own (200 MiB for rclone, 8 MB for the AWS CLI) and get a 405 when they do. Raise theirs to match ours: `--s3-upload-cutoff 2G` for rclone, `s3.multipart_threshold = 2GB` in `~/.aws/config` for the AWS CLI. `s3cmd` needs that hostname too.

---

## Two kinds of zone

A zone is the unit of URL, configuration and statistics. Its name is the first path segment: `https://cdn.example.com/assets/images/logo.png`. Give it a **hostname of its own** on the settings page (point a CNAME at the service first) and the same file is `https://assets.acme.com/images/logo.png`, with nobody else's name in the URL; the original `cdn.example.com` spelling keeps working too.

| | **Storage zone** | **Pull zone** |
| --- | --- | --- |
| Where files come from | you upload them | mirrored from an origin you name |
| Kept in | S3, plus the local cache | the local cache only |
| Good for | releases, images, packages, video | putting a cache in front of a site you already run |

A pull zone fetches a path the first time someone asks for it and keeps the answer for as long as the origin's own `Cache-Control` allows, clamped to limits you set. It collapses concurrent misses into a single origin request, remembers "not found" briefly, never caches an origin error, and **keeps serving its copy when the origin goes down**, which is the reason to put one in front of a site at all.

---

## Keeping things private

Zones are public by default. Make one private and its files are served only through a link this service signed:

```bash
cdn url reports 2026/q3.pdf --expires 2h
# https://cdn.example.com/reports/2026/q3.pdf?exp=1789…&sig=…
```

The browser has a **Sign** button next to every file that does the same. Each zone has its own signing key, so a leaked link opens that zone and nothing else; a link to one file does not open its neighbours.

Public zones can also carry a **referrer allowlist** (so other sites cannot embed your images and bill you for the traffic) and a **CORS origin list**, for fonts and modules loaded by script. Both are on the zone's settings page.

---

## Knowing what is being downloaded

Every delivered response is counted, per file, per hour, without touching the database on the request path. The zone list shows a week of traffic; the file browser shows a month of downloads per file. Each zone's **Statistics** page shows a daily chart, the cache hit ratio, the busiest paths, and what was asked for that could not be served, with a CSV export. The same numbers come back as JSON:

```bash
curl -H "Authorization: Bearer $CDN_TOKEN" \
  "https://cdn.example.com/_/api/v1/zones/assets/stats?days=30"
```

The counted bytes are the ones that actually reached the socket: a range request counts its range; an abandoned download counts what it transferred.

---

## Narrowing what a key can do

A token carries scopes (`read`, `write`, `delete`, `purge`, `stats`), and can be pinned to one zone and to one path prefix. A deploy key for `releases/` cannot replace the site's index page and cannot list beyond its prefix, since a listing that ignored the restriction would hand over every path in the zone anyway:

```bash
cdn create token "release uploads" --org acme --zone acme-assets \
    --prefix releases/ --scopes write
```

The same two fields are on the token form in the browser.

---

## Keeping a lid on it

A zone can carry a **storage quota**: set it on its settings page, which also shows what the zone and the whole organisation are holding. An upload that would go over is refused before the bytes are transferred, with `507`, or `QuotaExceeded` for an S3 client. Nothing is ever deleted to make room.

Whole-organisation caps belong to the operator, so they are set from the command line:

```bash
cdn set quota org acme 100G
cdn set quota zone acme-video 20G
cdn set quota org acme none        # lift it again
```

---

## Knowing who changed what

Every write is recorded: files uploaded, replaced and deleted, caches purged, zones created and reconfigured, tokens minted and revoked, people invited and disabled. Owners read it on the **Audit** page, which says who did it, from which address, and for a settings change what the value was before. Downloads are not in there; the statistics page has them.

It is kept for a year and pruned by the same worker that folds the statistics.

---

## Deploying it

The deployment is one [Hop3](https://hop3.cloud) app, described by [`hop3.toml`](hop3.toml): a web process, a PostgreSQL addon, a persistent cache directory, and two small workers. One evicts the least recently read files when the cache passes its high-water mark; the other folds old statistics into daily rows.

```bash
cdn migrate                    # runs before the new code serves traffic
cdn sweep  --interval 300      # the cache sweeper
cdn rollup --interval 3600     # statistics rollup and audit retention
cdn reconcile                  # compare the index against the object store
cdn reconcile assets --repair  # …and fix what disagrees
```

Run `reconcile` as a nightly cron: an upload writes the bytes before the index row and a delete removes them after it, so a crash in between leaves a blob nothing points at or a row pointing at nothing. It reports by default, exits non-zero when something disagrees, and only changes anything with `--repair`.

Configuration is entirely by environment variable; a missing one stops the process at boot, before the first request that needs it. The variables are `DATABASE_URL`, `CDN_BASE_URL`, `CDN_SECRET_KEY`, `CDN_CACHE_DIR`, `S3_ENDPOINT_URL`, `S3_BUCKET`, `S3_ACCESS_KEY_ID`, `S3_SECRET_ACCESS_KEY`, and optionally `S3_REGION`, `CDN_MAX_UPLOAD_BYTES`, `CDN_CACHE_HIGH_WATER`, `CDN_CACHE_LOW_WATER`, `CDN_S3_HOST`. Sizes are written the way a proxy writes them: `2G`, `500M`, `4K`.

`CDN_MAX_UPLOAD_BYTES` has to stay under whatever proxy sits in front, or that proxy refuses a large upload before the application can say anything useful about it. On Hop3 the app declares its own limit: `[proxy] max-body-size` and `request-buffering` in `hop3.toml`, the latter so a 2 GB upload does not need 2 GB of scratch space on the proxy. Anywhere else it is nginx's `client_max_body_size` and `proxy_request_buffering`, set by the operator.

One token may make 1200 writes a minute, a full minute's worth allowed to arrive at once: enough that a deploy pushing a thousand files never notices and little enough that a runaway script does. Set `CDN_WRITE_LIMIT_PER_MINUTE` to change it, or to `0` if your gateway already does this. Reads are not limited here; that belongs in nginx, where it costs nothing per request.

---

## Documentation

The full documentation lives in `docs/src/`, built with [Zensical](https://zensical.org): getting started, guides for each way in, the reference for the CLI and both APIs, and the architecture.

```bash
make docs         # into docs/site/
make docs-serve   # a live-reloading preview
```

---

## Developer notes

**Status: M7.** It includes storage and pull zones, public or private, with tokens, the web UI, signed URLs, statistics, purge, an S3-compatible endpoint, custom domains, an audit log, storage quotas, reconciliation and per-token write limits. The functional specification is [`notes/01-specs.md`](notes/01-specs.md); read it before changing anything structural.

**Stack:** Litestar, SQLAlchemy 2 with Advanced Alchemy, PostgreSQL, aiobotocore, Jinja. HTML is server-rendered, with one stylesheet and no build step.

```
src/abilian_cdn/
  app.py        Litestar factory, middleware, lifespans
  db.py         database wiring, sessions outside a request
  models.py     the schema
  delivery.py   the read path: validators, ranges, CORS, cache status
  origin.py     pull zones: fetch, TTLs, single flight, stale-if-error
  blobs.py      the order S3 and the local cache are touched in
  cache.py      the local disk cache and its sweeper
  api.py        the token API           web.py     the browser UI
  s3api.py      the S3 face             sigv4.py   the signatures it checks
  stats.py      counting and reporting  client.py + cli.py   the `cdn` command
```

**Everything the service exposes for itself lives under `/_/`**: `/_/api/v1` for the token API, `/_/ui` for the browser, `/_/health` for the load balancer, `/_/schema` for OpenAPI. The rest of the root namespace belongs to zone names.

**Tests**: three tiers, 642 of them:

```bash
make services              # createdb abilian_cdn_test + Garage
make test                  # everything
uv run pytest -m unit      # no services needed
make test-cov              # coverage
make lint                  # ruff, ty, pyrefly, mypy
```

`a_unit` needs nothing, `b_integration` talks to a real PostgreSQL and a real Garage, and `c_e2e` starts a uvicorn and drives the `cdn` command over HTTP the way CI would. A tier whose services are missing skips with a message saying what to start. **One run at a time per database:** the suite empties the tables when it starts, so two concurrent runs sharing `abilian_cdn_test` wipe each other and fail in ways that look like product bugs. Point a second run elsewhere with `CDN_TEST_DATABASE_URL`. CI does not run the service-backed tiers yet; there are no service containers in the workflow.

**Migrations** live in `src/abilian_cdn/migrations`, so they ship with the wheel:

```bash
make migrate                      # bring a database up to head
make migration m="what changed"   # generate one after changing models.py
```

Read the generated file before committing it: autogenerate proposes `NOT NULL` columns with no backfill, which fails on any table that already has rows. A test fails if the models and the migrations disagree.
