Heddle + Execution Host research adopter
Research in progress · not publicly availableCan one assistant stay useful across long-running projects?
The research question is whether one durable assistant can stay with an individual across user-created projects, preserve the work that matters, absorb authorized change, choose a useful next step, and remain within real product authority without a workflow written for every scenario.
Current status · September 2026
The product direction is accepted and the generic foundation is being implemented in Heddle and the private Execution Host research environment. The Assistant itself is not a shipped service, public beta, or Heddle feature. No claim of generally useful autonomous assistance has been earned yet.
The hypothesis
Given an intelligent model, stable assistant identity, durable free-form working files, separate long-term memory, authorized observations of what is new, bounded tools, and a reliable wake lifecycle, can the assistant notice what would help its user, consider consequences, and take a useful action without the product encoding a workflow for that particular domain or scenario?
What the experiment must prove
A convincing result requires continuity, judgment, authority, and cross-domain transfer—not a scripted demo that follows a predetermined checklist.
- Resume an active focus after a cold Runtime or process replacement.
- Reconcile new facts and constraints with existing analysis instead of restarting from a fresh prompt.
- Keep temporary plans, drafts, comparisons, and open questions in a durable working set while reserving long-term memory for stable user knowledge.
- Choose among continuing, reprioritizing, deferring, asking for approval, completing, or doing nothing.
- Evaluate consequence, confidence, and authority before an external or difficult-to-reverse action.
- Use the same workspace, wake, memory, and capability machinery in at least two unrelated domains.
- Leave a clear handoff so the next wake can continue coherent work.
- Return an honest no-action outcome when nothing helpful should happen.
Why home renovation is the first pressure test
Renovation is useful because it lasts for months and scatters context across messages, floor plans, renderings, photos, quotes, product links, measurements, meetings, private preferences, and decisions that later change. It creates the continuity, provenance, privacy, and consequence pressures a standing assistant should handle.
The product is not an AI design service and is not intended to create a new queue of speculative ideas for an architect or interior designer. The first assistant serves the individual owner: it organizes sources, preserves decision rationale, compares options, prepares private analysis, and may draft communication. Externally visible commitments remain product-controlled and approval-bound.
Renovation is therefore a dogfood and evaluation environment, not the permanent product boundary. The same generic foundation must later work in another unrelated project domain without gaining renovation-specific workflow code.
The assistant relationship, runtime, and product world remain separate
The Assistant is an adopter product. Heddle and the Execution Host provide reusable execution, but they do not own the user's projects, sources, permissions, or external truth.
The Assistant product
Owns the durable relationship with the user and the product semantics around projects, sources, permissions, decisions, approvals, and effects.
- One assistant relationship across user-created projects
- Project, Connection, Source Event, provenance, and visibility
- LINE as the first connector—not the product boundary
- Canonical product records, approvals, external effects, and UI
Heddle runtime
Owns the model/tool loop and the generic mechanics a product should not rebuild for every assistant experience.
- Conversation, tools, approvals, activity, traces, and artifacts
- Heartbeat execution and bounded autonomous cycles
- Separate working-set and long-term-memory semantics
- No renovation, LINE, or The Assistant domain vocabulary
Execution Host
Owns the hosted scope and lifecycle needed to restore the same assistant workspace on a replaceable Runtime and settle execution truthfully.
- Verified identity and no-widening scope binding
- Restore before tools become available
- Checkpoint before terminal success
- Isolation, cancellation, recovery, and storage generations
The first behavioral proof
The foundation should be judged by observed behavior across process replacement and unrelated scenarios, not by architecture diagrams alone.
- Wake 1 creates a recognizable working artifact and records the current focus.
- The Runtime is replaced; Wake 2 recovers the exact committed working state under the same verified identity.
- New authorized facts change the existing analysis rather than causing a fresh-start answer.
- One reversible action completes, one consequential action stops for approval, and one wake chooses no action.
- The same harness repeats the behavior in a second unrelated domain.
- Every attempted model, tool, and external effect is accounted for, including retries and failures.
What is being built first
- A bounded, portable `working/` directory separate from long-term `memory/`.
- Verified four-field scope binding so conversation and background work share files only for the same assistant context.
- Durable object-storage generations with restore-before-tools and checkpoint-before-success ordering.
- Cold and warm recovery tests that prevent a failed invocation from leaking dirty local state into the next one.
- A later adopter product boundary for connections, event horizons, approvals, effects, and user-facing truth.
Follow the evidence, not the Jarvis story
The goal is a falsifiable product and systems experiment: can stronger model intelligence become useful delegated work when it receives durable context, exact authority, and a reliable place to continue?