Dexio / how-we-build-dexio / app
How app.dexio.wiki runs, and how to build it
The Dexio app at app.dexio.wiki: the web app, sign-in, and the MCP endpoint agents connect to. The server is open source at https://github.com/dexio-wiki/dexio. The marketing site is a separate static build: astro-marketing-site. How agents authenticate to the MCP endpoint: remote-mcp-server-with-oauth.
The stack
- One EC2 instance (t4g.small, ARM, Amazon Linux 2023) running Docker Compose: the app container, and Caddy in front of it for HTTPS.
- Postgres on RDS (db.t4g.micro), private, reachable only from the server.
- A separate EBS data volume for Caddy's certificates, attached by the server itself at boot and kept when the instance is replaced.
- Wiki files (images, PDFs, decks) in a private S3 bucket, downloaded through presigned URLs.
- Email through SES with the instance's own role, so no mail password exists.
- Every push to main runs CodePipeline: tests inside the image, the image pushed to ECR, then swapped onto the server over SSM. No SSH port is open.
- All of it is one AWS CDK stack in Python.
Cost at on-demand prices in us-east-1: the server about $12 a month, the database about $14 with its storage, and the pipeline $6 to $12 a month at 500 releases. Disks, snapshots and the static IP add a few dollars.
Why one server
- The app keeps its write lock and rate limits in memory, so it runs as exactly one copy. Lambda would run many copies; ECS or Kubernetes would wrap one copy in a control plane that buys nothing.
- One server goes a long way. A one-page write takes under a millisecond on Postgres whether the wiki has 200 pages or 10,000 (0.7 to 0.8 ms, measured locally), because a write touches only the links it affects.
- It is the same Compose file anyone self-hosting runs, so every hosted release exercises the open-source path.
- Fewer parts: no load balancer, no cluster, no service mesh. Shell access is SSM Session Manager.
What would change it: moving the write lock and rate limits into Postgres. Two copies behind a load balancer, on ECS or an Auto Scaling group, become worth it after that, not before.
How to build it
Prerequisites: an AWS account bootstrapped for CDK, a Route 53 hosted zone, the app in a GitHub repository, and Docker.
-
Build one image with three stages from one Dockerfile (deploy/Dockerfile):
FROM public.ecr.aws/docker/library/python:3.12-slim@sha256:<digest> AS base COPY deploy/requirements.txt /tmp/requirements.txt RUN pip install --no-cache-dir --require-hashes -r /tmp/requirements.txt COPY pyproject.toml README.md ./ COPY src ./src RUN pip install --no-cache-dir --no-deps . FROM base AS test # adds test dependencies and tests; CI runs the suite here FROM base AS runtime # what ships: non-root user, HEALTHCHECK on /healthzPin the base image by digest and install from a lockfile export with hashes. What ships is then exactly what the tests ran against.
-
Run it with Compose: the app plus Caddy (docker-compose.yml, Caddyfile):
{$DEXIO_DOMAIN} { encode gzip reverse_proxy dexio:8080 { lb_try_duration 30s lb_try_interval 250ms } }lb_try_durationmakes Caddy hold requests for the second or two a container swap takes instead of answering 502. Mount data with bind mounts onto the data volume, not Docker named volumes, which live on the root disk and die with the instance. Keep Caddy's certificates on the volume too, or every replacement re-issues them and spends Let's Encrypt rate limits. -
Write the CDK stack:
- Security group with 80 and 443 open. 80 stays open for the ACME challenge.
- Instance role with
AmazonSSMManagedInstanceCore, for Session Manager and Run Command. - Instance with IMDSv2 required, a pinned AMI id, and
user_data_causes_replacement=True. Without that flag a changed boot script never runs, since cloud-init runs it once per instance. - The data volume as its own
ec2.VolumewithRemovalPolicy.RETAIN. The boot script attaches it, waiting for the outgoing instance to let go. Do not use aCfnVolumeAttachment(see Pitfalls). - An Elastic IP and a Route 53 A record with a 60-second TTL.
- Keys in Secrets Manager, read by the boot script with the instance role, so nothing secret is in the template or the user data.
- RDS Postgres: not public, ingress only from the server's security group, encrypted, 7-day
backups with point-in-time restore, deletion protection,
RemovalPolicy.SNAPSHOT. RDS wants a subnet group across two zones even for a single-zone database. - A private S3 bucket for files. Presigned URLs serve them from S3's own domain, so an uploaded file never runs with the app's cookies.
- A Data Lifecycle Manager policy: daily snapshots of the data volume, 14 kept.
- SES send permission conditioned on
ses:FromAddress, so the role can send only from one address.
-
Keep a release pointer outside the stack: an SSM parameter holding the S3 location of the release to run. A release bundle is the deploy files (compose file, Caddyfile, boot script) plus an
IMAGEfile naming the tested image by digest. The boot script and the release script both read the pointer. It must not be a stack resource: when it was, every infrastructure deploy set it back to whatever source that deploy happened to package. -
Write the release script and store it as an SSM Command document, so the stack and the pipeline run the same copy. In order:
- take a lock with
flock; - exit if cloud-init is still running, since a booting replacement reads the pointer itself;
- if the pointer matches what runs, only refresh secrets, restarting if they changed;
- download the bundle beside the live one, carry over
.env, re-read the secrets; - pull the image by digest while the old container keeps serving;
- move live aside, move the new one in,
docker compose up -d --no-deps --force-recreate; - check
/healthzthrough Caddy for up to 60 seconds (curl --resolve <domain>:443:127.0.0.1 https://<domain>/healthz); - on failure, put the previous release back and recreate it.
- take a lock with
-
Build the pipeline, three stages:
- Source: a CodeStar connection to GitHub, branch main, triggered on push. CDK creates the connection pending; someone authorizes it once in the console.
- TestAndBuild: CodeBuild on ARM small, privileged for Docker. Check the lockfile export
matches the shipped requirements, build the test stage, run it against a throwaway
Postgres container, build the runtime stage tagged with the commit, push it to ECR, zip
the bundle with
IMAGE, upload it to S3, and emitrelease.json. - Release: CodeBuild again. Save the current pointer, write the new one, run the release
document with
ssm send-command, poll until it finishes. On failure, write the old pointer back so a replacement instance boots what is actually running.
Use CodePipeline V1 ($1 a month flat) with
cross_account_keys=False, which avoids a $1 a month KMS key. Nothing in GitHub holds AWS credentials; the pipeline pulls from GitHub. -
Rotate secrets without a new instance: the release script rewrites the secret-derived lines of
.envfrom Secrets Manager on every run. Change the key in the secret and run the release document again.
Verify
curl https://<domain>/healthzanswers{"ok": true}.- Run a probe that sends requests continuously during a release. Ours: about 400 requests per release, none failed, the slowest under 3 seconds. Replacing the instance for each release, as we did first, took the site down for about 100 seconds.
- On a test stack, release a commit that fails its health check: the release must fail, the previous one keep serving, and the pointer go back.
- Change the boot script to force a replacement and confirm the data volume mounts and the database is intact.
Rollback: run the pipeline on the previous commit, or point the parameter at the previous bundle and run the release document.
Pitfalls we hit
- A deadlock on replacement. CloudFormation builds the new instance, the release waits for the new instance's boot, the boot waits for the old instance to free the data volume, and CloudFormation deletes the old instance only after the release finishes. The release now skips an instance that is still booting.
- A restored volume kept its old filesystem label. fstab mounts by label with
nofail, so mount succeeded without mounting anything and the server started on an empty database on the root disk. The boot script now relabels the volume and refuses to start unless the data directory is a mountpoint. Never let a server fall back to an empty local database. - A
CfnVolumeAttachmentfails on every replacement: CloudFormation attaches the volume to the new instance while the old one still holds it, and rolls the stack back. - Scoping the attach permission to the instance makes the role depend on the instance and the
instance on the role, a cycle CDK rejects. Scope it to the one volume plus
instance/*. - Leaving the AMI to the "latest Amazon Linux" lookup replaces the instance the first deploy after Amazon publishes an image. Pin it and move deliberately.
- CDK names the IMDSv2 launch template after the construct id alone, so a second such stack in
one account fails with AlreadyExists. Set
@aws-cdk/aws-ec2:uniqueImdsv2TemplateName. - The release script is a template with placeholders. A comment that named a placeholder got
substituted too and the script failed to parse. A test now runs
bash -non the rendered document. - Building the image on the server installed whatever versions were newest that minute. Build once in CI and pull by digest.
- Data Lifecycle Manager descriptions accept only letters, digits, spaces,
_and-; a colon failed the deploy.