A Python backend for AI data pipelines

Astrocoda is boilerplate for the part everyone underestimates: getting text in, handing the work to a queue, and getting typed JSON and vectors back out. docker compose up starts five services. The request handler returns 202 in about a millisecond.

Read the source
46 hermetic tests
202 on trigger
0 license servers
MIT licensed

Request path

Six steps, all of them already written

Nothing here is a design document. Each line below is a function or route that exists in the repository today, with the behaviour described as implemented rather than as intended.

  1. POST /v1/pipelines/trigger

    Validates the payload, writes a run row, enqueues an ARQ job and returns 202 with a job id. No LLM call happens in the request path, so a slow provider cannot add latency to your API.

  2. chunk_text()

    Splits input into windows of at most CHUNK_SIZE characters, cutting at the last whitespace in the final 40% of the window so words are not broken in half while forward progress is still guaranteed.

  3. embed_batch()

    Sends chunks in batches of EMBEDDING_BATCH_SIZE and re-sorts the response by index, because embedding APIs do not guarantee response order.

  4. ExtractedInsight

    A Pydantic model with a summary of 1 to 2000 characters, a sentiment constrained to three values, and 3 to 8 keywords. Instructor retries on validation failure instead of returning malformed JSON.

  5. persist

    Vectors to Qdrant, structured rows to PostgreSQL. A unique index makes duplicate chunks impossible, and a job resumes at the first missing chunk rather than starting over.

  6. GET /v1/pipelines/jobs/{id}

    Reports queued, running, completed or failed, with per-chunk progress and a duration.

Interactive

Watch a run

Paste text and it will be chunked, embedded, extracted and validated using the same code paths as the backend.

Simulation. This runs entirely in your browser. No LLM is called and nothing is sent anywhere. Chunking and schema validation are the real algorithms; the extraction step is deterministic stand-in logic, so the output below is illustrative, not model output.

Try a longer paste to see the chunk count change. CHUNK_SIZE is shown as the real default from .env.example.

Idle.

Distribution

The download is a check, not a copy

This is the part that is genuinely unusual, so here is exactly what happens. The release zip contains astrocoda.manifest.json, an Ed25519 signature over a sorted map of every shipped file to its SHA-256.

$ astrocoda init --template ./astrocoda-0.1.0.zip
manifest signature   valid
schema version       1
files verified       41
scaffolded           ./my-project

init refuses to scaffold when there is no manifest, when the signature does not verify, when a listed file is missing, or when a hash differs. It also copies only the paths the manifest lists, so an extra file dropped into a template does not reach your project.

The public key ships with the CLI, so the check runs offline. There is no server that could be pressured into vouching for a download, and no server that can be switched off.

Practical consequence: a mirror of the zip is as trustworthy as the original, which is what makes a single download link survive being reposted.

Contents

What is in the box

Runtime stack
Layer Choice
API FastAPI, async end to end
Queue ARQ on Redis, with retries and timeouts
Rows PostgreSQL via asyncpg and SQLAlchemy
Vectors Qdrant
Extraction Instructor against any OpenAI-compatible endpoint
Payments Stripe, webhook signature verification
app/
API routes, worker, database and vector access
cli/
Signed-manifest verification and release download
tests/
46 tests, no services required
Dockerfile
One image, API and worker entrypoint
docker-compose.yml
Five services, wired
.github/
CI across Python 3.11 to 3.13 plus a Docker build

Pricing

It is free, and there is nothing to unlock.

Astrocoda is MIT licensed and the whole repository is public. Use it, fork it, put it in production, change it. There is no licence key, no account and no upgrade tier, because the CLI does not check any of those before handing you the code.

Free forever, no account

  • Everything in the public repository, including the CLI
  • Signed releases you can verify offline against a manifest
  • The supply-chain check that ships with astrocoda init
  • Issues and fixes on the public tracker
Get the code

astrocoda init asks once whether you want update news by email. Say no and nothing is sent — there is no list entry to leak.

FAQ

Questions worth answering plainly

Why is the repository public if you sell it?

Because a licence file that hides the code is a worse product. The MIT grant is real and permanent, so pretending otherwise would only cost trust. The paid part is provenance, updates and support.

Do I need the paid version to use it?

No. Clone it, run docker compose up, and you have the whole system including the signed-manifest tooling. The purchase is a convenience and a support contract.

What model providers does it work with?

Anything OpenAI-compatible. Set OPENAI_BASE_URL and the key for that provider, and the chat and embedding models follow. It ships configured for OpenAI and was developed against OpenRouter.

Is there a license key?

No. The CLI is free, MIT licensed and completely ungated — there is no account, no activation call and nothing to log in to. You can also scaffold straight from the public repository if you would rather not install anything at all.

What happens if I fork it?

Your fork works immediately. It will not verify against the seller's manifest, because the signature covers specific bytes, and the release tooling is not published. That is the intended line: MIT covers the application code, not the signing key.