Stop supervising tasks. Start delegating objectives.

Coresource AI
announcementcoresourcelong-horizon-agents
Stop supervising tasks. Start delegating objectives.

Introducing Coresource — built to engineer for days, not minutes.


We recently pointed Coresource at one hard objective and let it run.

When it stopped, it had processed ~1.5 billion tokens of context, worked for ~50 hours straight, and shipped ~50,000 lines of working code — and it ran to completion. One run. One agent. No human in the loop driving it step by step.

That run is why Coresource AI exists, and it is the clearest picture we have of where software engineering is going.


The agent era has a ceiling, and everyone keeps hitting it

The last two years gave us a generation of remarkable coding assistants. They autocomplete, they answer, they refactor a function, they draft a pull request. Used well, they are a real multiplier.

But watch any of them on a genuinely large task and the same thing happens: they stall. They lose the thread after a few minutes and a few thousand tokens. They forget what they decided ten steps ago. They need a human to re-aim them constantly, to remember the plan, to notice when the plan was wrong, to pick the work back up after every interruption.

That is not a model problem. The models are extraordinary. It is a horizon problem — the length of coherent, self-directed work a system can sustain before a human has to step back in.

Real software engineering does not happen inside a few-minute horizon. It happens across hours and days: investigate, plan, build, hit a wall, re-plan, verify, keep going. The entire hard part of engineering lives in the part of the timeline that today's assistants cannot reach.

Short-horizon autonomy is a feature. Long-horizon autonomy is a category. We are here to build the category.


What Coresource is

Coresource is a long-horizon autonomous software-engineering system. You give it an objective, not a prompt. It owns the whole arc:

  • Plan — it decomposes a real objective into milestones, steps, and execution waves.
  • Execute — it does the work durably, so a long run survives interruptions, failures, and retries instead of collapsing.
  • Re-plan — when reality diverges from the plan, it steers itself: it revises the plan mid-flight rather than waiting for a human to notice.
  • Verify — it checks its own work against the objective and keeps going until the objective is actually met.

The difference is not that Coresource is a smarter autocomplete. The difference is that Coresource stays coherent across a horizon long enough to finish serious work — and it steers itself when the path changes, which on any real task it always does.


The proof: one run, by the numbers

We do not want you to take "long-horizon" on faith. Here is a single, real Coresource executor run — the one that convinced us we had something worth building a company around:

What it didThe number
Context processed~1.5 billion tokens (1,506,242,499)
Wall-clock duration~50 hours (52.8h, to completion)
Code shipped~50,000 lines (113 commits · 131,217 insertions · 120 files)
Plan structure it managed8 milestones · 36 steps · 8 execution waves
Self-correction47 step attempts · 201 revision-ledger entries
Cost$1,461.48

The objective it was given was not a toy: implement steerability across the entire agent lifecycle — investigation, planning, and execution — and fold prior research into a live codebase. It planned that work, executed it, re-planned 201 times when it learned something new, and verified itself to a finished state.

One honest caveat Of those ~1.5 billion tokens, roughly 1.37 billion were cache-read input tokens and about 4.3 million were newly generated completion tokens. So the precise claim is 1.5 billion tokens processed — an agent that sustained and reasoned over that much context across 50 hours — not 1.5 billion tokens written. We will always tell you which number is which. The long-horizon story is remarkable enough without inflating it.


Why this is different

Today's tools optimize the first ten minutes of a task. Coresource optimizes the last ten hours — the part where the plan breaks, where work has to be picked back up, where a human normally has to take the wheel.

Two properties make that possible, and they are the things short-lived assistants structurally cannot do:

  1. Durability — a Coresource run is built to survive failure and resumption. A 50-hour task is not 50 hours of luck; it is an execution model designed to not fall over.
  2. Self-replanning at scale — Coresource revises its own plan as it learns. 201 revisions in a single run is not the system failing; it is the system steering. That is what a human engineer does all day, and it is what an assistant that needs you to re-aim it cannot.

Single-shot assistance asks: what's the next edit? Coresource asks: what is the objective, and what will it take to actually finish it?


Where this goes

We think the unit of software work is about to change. Not from "human writes code" to "AI writes code" — that framing is already old. The shift is from tasks you supervise to objectives you delegate: work measured in hours and days, planned and steered and verified by the agent itself, that you check the way you check a finished result rather than a half-built one.

That is engineering at machine scale — software that plans, builds, and steers itself toward a goal. Coresource is building the system that makes it real, and this is the first step across that horizon.


Get early access

Coresource is moving into the hands of a first set of teams now.

→ Sign in to Coresource.

We're also hiring the engineers and researchers who want to build the long-horizon era. If that's you, come build it with us.