A Python backend for AI data pipelines
Astrocoda is boilerplate for the part everyone underestimates: getting text
in, handing the work to a queue, and getting typed JSON and vectors back
out. docker compose up starts five services. The request
handler returns 202 in about a millisecond.
Request path
Six steps, all of them already written
Nothing here is a design document. Each line below is a function or route that exists in the repository today, with the behaviour described as implemented rather than as intended.
The request returns before the model is called. That is the whole point of the queue: provider latency never reaches your API.
-
POST /v1/pipelines/trigger
Validates the payload, writes a run row, enqueues an ARQ job and returns
202with a job id. No LLM call happens in the request path, so a slow provider cannot add latency to your API. -
chunk_text()
Splits input into windows of at most
CHUNK_SIZEcharacters, cutting at the last whitespace in the final 40% of the window so words are not broken in half while forward progress is still guaranteed. -
embed_batch()
Sends chunks in batches of
EMBEDDING_BATCH_SIZEand re-sorts the response byindex, because embedding APIs do not guarantee response order. -
ExtractedInsight
A Pydantic model with a
summaryof 1 to 2000 characters, a sentiment constrained to three values, and 3 to 8 keywords. Instructor retries on validation failure instead of returning malformed JSON. -
persist
Vectors to Qdrant, structured rows to PostgreSQL. A unique index makes duplicate chunks impossible, and a job resumes at the first missing chunk rather than starting over.
-
GET /v1/pipelines/jobs/{id}
Reports
queued,running,completedorfailed, with per-chunk progress and a duration.
Interactive
Watch a run
Paste text and it will be chunked, embedded, extracted and validated using the same code paths as the backend.
Simulation. This runs entirely in your browser. No LLM is called and nothing is sent anywhere. Chunking and schema validation are the real algorithms; the extraction step is deterministic stand-in logic, so the output below is illustrative, not model output.
Try a longer paste to see the chunk count change. CHUNK_SIZE is
shown as the real default from .env.example.
Idle.
Distribution
The download is a check, not a copy
This is the part that is genuinely unusual, so here is exactly what happens.
The release zip contains astrocoda.manifest.json, an Ed25519
signature over a sorted map of every shipped file to its SHA-256.
$ astrocoda init --template ./astrocoda-0.1.0.zip
manifest signature valid
schema version 1
files verified 41
scaffolded ./my-project
init refuses to scaffold when there is no manifest, when the
signature does not verify, when a listed file is missing, or when a hash
differs. It also copies only the paths the manifest lists, so an extra
file dropped into a template does not reach your project.
The public key ships with the CLI, so the check runs offline. There is no server that could be pressured into vouching for a download, and no server that can be switched off.
Practical consequence: a mirror of the zip is as trustworthy as the original, which is what makes a single download link survive being reposted.
Contents
What is in the box
| Layer | Choice |
|---|---|
| API | FastAPI, async end to end |
| Queue | ARQ on Redis, with retries and timeouts |
| Rows | PostgreSQL via asyncpg and SQLAlchemy |
| Vectors | Qdrant |
| Extraction | Instructor against any OpenAI-compatible endpoint |
| Payments | Stripe, webhook signature verification |
- app/
- API routes, worker, database and vector access
- cli/
- Signed-manifest verification and release download
- tests/
- 46 tests, no services required
- Dockerfile
- One image, API and worker entrypoint
- docker-compose.yml
- Five services, wired
- .github/
- CI across Python 3.11 to 3.13 plus a Docker build
Pricing
It is free, and there is nothing to unlock.
Astrocoda is MIT licensed and the whole repository is public. Use it, fork it, put it in production, change it. There is no licence key, no account and no upgrade tier, because the CLI does not check any of those before handing you the code.
Free forever, no account
- Everything in the public repository, including the CLI
- Signed releases you can verify offline against a manifest
- The supply-chain check that ships with
astrocoda init - Issues and fixes on the public tracker
astrocoda init asks once whether you want update news by
email. Say no and nothing is sent — there is no list entry to leak.
FAQ
Questions worth answering plainly
Why is the repository public if you sell it?
Because a licence file that hides the code is a worse product. The MIT grant is real and permanent, so pretending otherwise would only cost trust. The paid part is provenance, updates and support.
Do I need the paid version to use it?
No. Clone it, run docker compose up, and you have the whole
system including the signed-manifest tooling. The purchase is a
convenience and a support contract.
What model providers does it work with?
Anything OpenAI-compatible. Set OPENAI_BASE_URL and the key
for that provider, and the chat and embedding models follow. It ships
configured for OpenAI and was developed against OpenRouter.
Is there a license key?
No. The CLI is free, MIT licensed and completely ungated — there is no account, no activation call and nothing to log in to. You can also scaffold straight from the public repository if you would rather not install anything at all.
What happens if I fork it?
Your fork works immediately. It will not verify against the seller's manifest, because the signature covers specific bytes, and the release tooling is not published. That is the intended line: MIT covers the application code, not the signing key.