Runtime and Control Plane
Extend manifests, harness behavior, workspaces, and agent-authored mutations without duplicating policy.
Agent Runtime Responsibilities
The agent app owns runtime mechanics: claims, backend calls, manifest loading, workspace hydration and sync, harness invocation, event normalization, heartbeats, and graceful shutdown. The backend owns policy and canonical transitions.
Depend on TAgentHarness, not Codex SDK types. Normalize provider events at the adapter boundary.
Context Manifest
The manifest contains runtime configuration, static MCP keys, organization integration IDs, missing required-integration markers, authorization epoch, context files, declared output directories, and optional scratch directories. It never contains provider credentials.
Context construction must page every filtered repository collection before selecting relevant dependencies, durability, human items, actions, task descendants, or integrations. Do not filter a single first page in memory: a valid item beyond that page would disappear from the run snapshot.
Task runs receive bounded summaries of recent attempts. Historical output inventories and bytes
are read on demand through oblivectl, including earlier attempts beyond the startup summary. New
manifests do not hydrate history beneath attempts/; retained legacy manifests remain readable.
Trajectories, logs, workspace locations, and raw provider payloads are not shared evidence.
Treat every manifest loaded from object storage as untrusted persisted input. Validate paths, expected sizes, hashes, roles, and read-only/output contracts before hydration.
When a task run starts, copy every selected mutable organization Context file to run-owned object storage before writing the manifest. Instructions and skills from a pinned agent-pack revision, retained prior-attempt output, and other already-immutable sources may remain referenced in place. This preserves the exact mutable boot bytes without duplicating immutable content.
The owner-facing run-context API is a separate allowlisted projection of the manifest. It returns only target path, display name, role, detected media type, size, and availability, then serves an exact recorded path through the authenticated content boundary. The runtime snapshot is redacted before delivery. Never expose source object keys, hashes, prompts, runtime or tool configuration, integration grants, output directories, leases, or operational locations. Legacy canonical Context may be served only when its current bytes match the manifest’s recorded SHA-256; otherwise retain the entry as unavailable.
Task workers pass a 100 changed-file ceiling to terminal workspace sync. The workspace manager must inventory and validate the full candidate set before its first object-store write. Chat sync does not use this task-only ceiling.
Every non-verification task manifest declares work/ as disposable scratch. Repository checkouts,
dependency trees, caches, tests, builds, and nested development symlinks may live there. The
workspace manager verifies the scratch root itself remains a real directory but never inventories,
uploads, receipts, hydrates, or counts its contents. outputs/ remains the only durable write
boundary; durable symlinks and changes outside declared boundaries fail closed.
When source work must pause for a human or action approval, Task Runtime writes a bounded patch and
continuation note beneath outputs/handoff/. The next attempt discovers and reads those exact files through oblivectl, creates a fresh
checkout in work/, and reapplies the patch.
This handoff is an agent instruction; the worker intentionally does not inspect scratch contents to
infer whether a patch was needed.
A run can finish while its task waits for a human response. That pause must leave downstream dependencies pending: success requires a successful task, and completion requires a terminal task. The answer resumes the original task; its eventual terminal result resolves the dependent work. Upgrades repair dependency links previously finalized from a nonterminal human interruption, preserving task identities, ownership, budgets, genuine failures and already completed work. Worker admission rechecks those canonical links even if the task still displays an older dependency blocker. Once all required links are satisfied, it starts the original queued request and clears resolved dependency blockers atomically. Human, integration, action and review waits still prevent execution; a duplicate wakeup cannot create another attempt.
Plans retain existing dependencies. Repeating the same edge and contract is idempotent; changing its condition or required flag is rejected atomically. Correctable plan validation errors use the existing bounded attempt retry, so the next run can inspect the error and correct the plan. Permission and exhausted-budget errors remain terminal; retries do not grant additional credits. When retrying a failed initiative, a cancelled review waiting on its coordinator is restored from its last committed amendment. The same reviewer, proposal, ownership and consumed credits remain.
Worker Leases
The backend owns both the heartbeat cadence and lease duration. A worker submits its identity and fencing token when it starts an attempt; the start response supplies the absolute lease deadline and heartbeat interval. Heartbeats renew that backend-issued deadline. Worker input cannot choose or extend the lease duration.
RUN_HEARTBEAT_INTERVAL_MS defaults to 60 seconds and RUN_LEASE_DURATION_MS defaults to 130
seconds on the backend; the lease must be longer than the heartbeat interval. Worker calls to the
backend are independently bounded by BACKEND_REQUEST_TIMEOUT_MS, which defaults to 30 seconds.
Failure to persist a heartbeat aborts local execution, while runbeat closes an attempt as soon as
its canonical PostgreSQL lease expires.
Run Activity Logging
Task-run activity is a safe operational projection, not a transcript or chain-of-thought record. The provider adapter exposes normalized events to a synchronous observer; the worker reduces them to categories, status, a short label, and an optional bounded detail. It never persists prompts, reasoning, assistant messages, command output, tool input/output, provider errors, credentials, or raw object-store locations. Known execution capabilities and integration secrets are redacted, workspace roots become relative paths, and database constraints enforce 160-character labels and 2,048-character details.
Each worker keeps one in-process queue for its active run: at most 256 unsent entries, one 20-entry
request in flight, a one-second flush cadence, a three-second request timeout, one bounded transient
retry, and a five-second terminal drain. Overflow retains recent activity, emits an omission marker,
and finalizes the run’s log state as partial. Delivery failures are rate-limited in process logs
and never fail or stall the agent run. There is no observer subprocess or log file to rotate.
Workers append directly to the fenced internal run API. PostgreSQL stores one run_logs row per
sequence and runs.log_state records not_recorded, recording, complete, or partial. This is
not an agent-authored mutation, so it does not go through oblivectl. Redis and Pub/Sub are not in
the path: durable HTTP batching is sufficient and avoids a second queue whose state would still
need reconciliation. runs.log_uri remains reserved for future large raw logs and is not used by
this projection.
The owner API serves tenant-scoped pages of at most 100 entries. The task log viewer uses the same
URL-backed state in its sheet and expanded presentations. It polls the latest page every three
seconds only while open, visible, and active; it replaces each bounded page, stops for terminal
runs, and never polls historical pages. SSE is intentionally reserved for continuously observed
chat: opening a run log occasionally does not justify a persistent connection or another fan-out
channel. Log rows cascade with the run, and pre-feature runs remain explicitly not_recorded.
Run Outcome Ownership
Successful execution has one writer. After the harness returns, the worker synchronizes declared
outputs, derives sanitized receipts, and sends the exact final response and uploaded artifact list
to run completion. The backend creates the immutable execution outcome and applies the task
transition in that same transaction.
Agents use oblivectl task outcome submit only for semantic branches such as a plan, review verdict,
verification report, promotion, human or action interrupt, deferred continuation, learning result, or classified failure.
Execution-scoped authorization rejects agent-authored execution outcomes. If a legacy or trusted
service-staged execution outcome already exists, completion still requires exact equality with the
worker-retained result.
Execution and research may submit deferred with a reason in summary, a future UTC resumeAt,
and a continuationPath beneath outputs/. Completion requires a nonempty file receipt and a
matching worker-synchronized artifact from this run. It finishes the attempt and returns the same
task to pending at the requested time, preserving owner, dependencies, phase and consumed credits.
The scheduler and worker admission both respect the due time. A task with no remaining run credit
fails explicitly. Pending actions and staged human questions require their own handoffs; deferral
cannot bypass approval, reconciliation, or a human decision. Discovery research retains its final
packet requirement when it completes, rather than on an unfinished deferred attempt.
A one-time learning schedule that encounters an active cycle retries after one minute. It remains enabled until it can create work, then disables after firing once. Recurring learning schedules advance their normal cadence when skipping an overlap.
An unstaged worker completion with unresolved actions becomes an action wait automatically.
Submission receipts cannot prove a completed write. Completion uses backend action results and
reuses authorized same-task outputs and command receipts from the last eleven retained runs,
preserving their original run IDs. Older ambiguous integration receipts are not proof of an
external effect. A schedule_enabled criterion checks the real schedule and optional frequency.
Coordinator repairs can append additionalAcceptanceCriteria without removing existing checks.
Queued actions treat an MCP isError result as a failed operation. HTTP 429 rejections and
transient failures before dispatch retry on the same action, at most three attempts, using a
persisted availableAt and the provider’s Retry-After when supplied. A transport failure after
dispatch keeps the outcome unknown for reconciliation; the runtime does not repeat a write
whose effect may already have happened.
Add Harness Behavior
- Extend the provider-neutral interface only for a real cross-adapter need.
- Keep Codex-specific configuration in the Codex adapter.
- Preserve trace, revision, lease, authorization, and cancellation semantics.
- Normalize events into existing agent runtime contracts.
- Add adapter and runtime tests.
- Update the developer guide when the public extension contract changes.
Add an oblivectl Command
oblivectl is the only authenticated state-changing/control-plane CLI available to agents. The
trusted chat/worker process keeps the global service token; the model process receives only a
short-lived capability bound to its organization, profile, and exact task run or chat turn.
- Confirm the workflow is a backend-owned mutation rather than a filesystem operation.
- Add a typed command parser in
packages/oblivectl. - Extend its backend client with the smallest request/response contract.
- Read organization, profile, task, run, and snapshot identity from hydrated context; read only the execution capability from the session environment.
- Add or update the backend service-authenticated endpoint.
- Preserve version, lease, ownership, and transition fencing.
- Produce complete observable output suitable for the calling skill.
- Update the
oblivectlagent skill and command reference. - Update backend OpenAPI documentation.
- Add CLI, client, route, and lifecycle tests.
Execution capabilities can read tasks and retained output files across departments in their active organization. Department executions mutate only their owned tasks in the active root. Operator executions coordinate within the active root. Chat turns may perform organization-level control-plane work only in response to the active user turn.
Codex event normalization forwards actual provider events only. Do not manufacture compaction events; persistent chat resumes its real harness session, while ephemeral task runs start fresh.
In local authentication mode, the harness adapter may recover one credential failure only when the deployment synchronizer installed a generation newer than the failed attempt and the provider stream emitted no item event. Harmless session/turn-start events are buffered across that retry. Any item event fences replay because a command, tool, message, or file effect may already have started. An unrecovered chat failure exposes an actionable, redacted authentication message; an unrecovered task attempt is non-retryable until an administrator updates the deployment credential.
Every new run receives a backend-built safe integration snapshot. Existing runs are not hot-patched. Material changes advance the organization authorization epoch, so a stale execution cannot continue through the integration gateway and must restart from the current snapshot. The snapshot reports control-plane readiness only and must not trigger a live provider health call.
A task-completion verify run is narrower than ordinary execution. Its manifest selects the
read-only sandbox, disables network and web search, supplies no output directories, MCP servers, or
integration IDs, and exposes only retained evidence. Backend authorization rejects task mutations,
human questions, schedules, Context promotion, and worker-owned execution outcomes from that run. It
may submit only its exact verification report; the backend owns merge, repair, human escalation, and
closure.
Do not add agent-facing wrappers for ordinary reads or declared-output writes when native filesystem access and deterministic validation are sufficient.
contextctl
contextctl verify is intentionally local and read-only. It validates staged structured Context. It
must not read files on behalf of the agent, write files, create manifests, call the network, or
promote canonical state.
Live integration context
contextctl integration list and contextctl integration show <key> read the current safe registry,
including contextRevision and operatingContext. Credentials are never returned.
contextctl integration update <key> --from-json - (also available through oblivectl) accepts:
{
"expectedRevision": 0,
"context": {
"organizationPurpose": "Source code hosting",
"departments": ["engineering"],
"resources": [
{
"kind": "repository",
"externalId": "org/product",
"name": "Product",
"url": "https://github.com/org/product",
"defaultBranch": "main",
"departments": ["engineering"]
}
]
}
}An update replaces the whole operating context. Read first and preserve fields you intend to keep. A stale revision returns 409; read the new context and reconcile before retrying. Only current Chat and Operator executions may save organization integration context. Saving does not start an agentic refresh. Connection and permission changes invalidate access independently of semantic Context.
Reading Retained Evidence
Use oblivectl task runs <task-ref> to page historical attempts, run show <run-id> for safe
run state and committed outcomes/checkpoints, and run files <run-id> to discover exact retained
outputs/ paths. run file read <run-id> <outputs/path> returns bounded UTF-8 or base64 content
with source task/run provenance. Follow list cursors and byte nextOffset; pass the returned etag
on subsequent chunks. Decode each chunk before joining bytes.
Same-organization evidence reads are already authorized. Agents read completed work and handoffs directly, without a human permission request, file relay, or a new task just to retrieve existing evidence. Task and chat prompts carry this rule on every invocation; catalog instructions route missing-input decisions through evidence discovery before asking for human judgment. An actual failed read is an access or availability failure, not missing human consent.
These GET routes remain available to read-only verification runs. They cannot change task ownership,
write files, access another organization, fetch raw trajectories or read arbitrary storage keys.
Failed terminal runs may contain useful partial outputs. Missing worker receipts on legacy files
are reported as unrecorded; only a full checksum match is verified. Changed content fails with
409 and unavailable source bytes with 404. Missing local hydration alone never means evidence was lost.
Failed worker runs retain receipts for successfully synchronized partial files. This preserves source checksums without treating partial work as successful execution.
Reusable Findings And Procedures
Agents search oblivectl knowledge search "topic" before repeating investigations and read a
relevant procedure with knowledge show <id>. Search returns compact findings, applicability,
versions, evidence references, and freshness; it does not hydrate earlier workspaces. The same
organization-scoped GET APIs are available to the owner and authenticated agent readers.
Successful task runs can optionally retain outputs/knowledge.json using the shared
@oblive/types/knowledge-contracts schema: up to five evidence-backed entries within 64 KiB.
Entries record a stable key, finding, applicability, procedure, audience, source versions,
validity of at most 90 days, and exact file hashes. A small PostgreSQL full-text index supports
search; the procedure remains in the immutable object-store packet. Reading verifies the packet
hash and rechecks current authority before returning content.
General organization findings are accessible across enabled profiles. Department visibility and
restricted connector grants are enforced independently. Source scope or credential changes hide
restricted records and mark general derived findings stale; expiry and supersession are explicit.
A current label means the recorded validity and source checks pass, not independent verification
of factual truth or the latest repository commit. Agents must check applicability before reuse.
No knowledge record grants integration access or action authority.
Completion queues publication atomically with worker receipts. Observation maintenance publishes at most two packets per tick, one per organization, with a ten-second storage deadline each. Storage outages retry the publication checkpoint without rerunning the task or external effects; invalid packets are rejected without reopening completed work. Scheduler and run recovery perform no knowledge object-store reads. Publication is version-fenced and newer source runs supersede older findings from the same profile even when retries finish out of order. Trashed tasks hide owned records, and shared citations preserve evidence needed by other work during permanent deletion.
Daily Learning And Suggestions
After setup, Operator and enabled departments learn every 24 hours. A cycle uses available organization context, reusable knowledge, work and fixed business metrics; it does not wait for four completed initiatives. Concurrent scheduled and on-demand root learning requests coalesce under a profile-scoped database lock, preserving the active task and its budget. Generation proposes up to five lessons and three ranked next steps with retained file evidence. A separate verification run checks the sources. The backend validates worker receipts and promotes only lessons scoring at least 95 without hard failures. Uncertain lessons are skipped without a human approval question. Missing metrics or sources limit conclusions rather than blocking the cycle.
Successful verification atomically replaces the profile’s suggestions, including with an empty list. Failed cycles keep the previous set. Suggestions contain a title, rationale, proposed owner prompt and evidence; they never execute work automatically. Delayed older cycles cannot overwrite a newer cycle’s result. Stable-key corrections replace prior lessons; the active working set keeps at most 24 lessons within 8 KiB, while immutable outcomes retain history.
Department runs receive compact organization and own-profile learning. Chat receives refreshed organization learning on every turn, including resumed sessions. Other departments’ complete memories and earlier workspaces are not automatically loaded; relevant retained knowledge remains available on demand.
oblivectl learning list and learning show <profile> provide scoped conclusions without loading
entire workspaces. Operator and Chat can read all enabled profiles; departments read organization
and own-profile learning. learning refresh <profile> returns the active learning task or creates
a pending task for the scheduler. Departments may refresh only themselves. Verification and stale
executions cannot refresh. Owner and internal HTTP routes use the same learning contracts.
Knowledge publication, analytical maintenance and metrics collection have independent poll guards. A failed or stalled collection cannot stop publication on later ticks; a stage does not overlap itself while it is still running.
Generation may propose up to three essential business-input questions. Verification must confirm
both necessity and retained evidence before the backend opens an Inbox question with
purpose: learning_input. A profile has at most three open learning questions, deduplicated by
stable key. These questions never block the completed learning task. Answering records a durable
refresh request; the scheduler waits for any active root learning cycle to finish, then creates a
fresh cycle with canonical, same-profile answers in learningInputs. Duplicate answers are
idempotent; a different response conflicts. Answers arriving during a cycle are retained for a
later cycle rather than added to its old snapshot. A lesson citing humanInboxItemIds must select
those answered inputs and independently verify their provenance before publication.
Human Communication
Chat exposes structured_questions.ask for essential operational decisions and review_context
for explicit Context-note approval. MCP, backend and frontend share the structured question schema
in packages/types. New operational requests contain at most three questions, permit deferral and
allow an owner-supplied alternative to selection options. The durable tool input identifies the
question; an answer uses the assistant-message/part reference, preventing collisions when provider
item IDs repeat. General chats permit deferral of older requests too. Skipping never implies
permission to execute dependent work. Context approval retains its separate validator.
Agent inbox requests can provide a structured presentation: required title and summary,
optional businessImpact and recommendation, and two or three options with stable IDs and
human-readable labels. An optional recommendedOptionId names one of those options. Owners can
choose an option, add context, or write their own response; recommendations are never submitted
automatically. Notifications normally need no options.
diagnostics is a separate bounded array of safe code/detail pairs. The inbox shows business
content first and keeps diagnostics in a collapsed Technical details view. Legacy message-only
records remain readable. The backend retains a compatible message derived from a presentation
and includes presentation and diagnostics in replay-conflict checks. A duplicate request does not
create a second question. Question staging, ownership, approval and task-resumption rules remain
unchanged.
Chat and Operator read available Context, knowledge, metrics and work before asking one to three material questions for a substantial initiative. Routine implementation choices and authorized cross-agent evidence reads require no new human consent. Support separates customer drafts from internal notes; Engineering verifies the promised delivery (a PR does not imply merge or deploy); Growth uses current brand context, measurable hypotheses and supported personalization.
Review Repairs
A child review’s amend verdict records a proposal against its exact source run and wakes the
nearest ancestor task or initiative coordinator. The reviewer waits with a typed review_repair
blocker; no human question is created from the review’s zero replan allowance.
The coordinator reads the proposal and submits a bounded amendment under its own identity and
budget. Its plan reopens both the repair targets and the addressed waiting reviewers, preserving
their IDs and owners. Reviewers resume in review after required dependencies finish. Child
re-execution consumes their remaining run credits, while graph amendments consume the coordinator’s
replan allowance. Pending proposals prevent coordinator approval. Exhaustion retains the existing
single human override; these changes do not reopen previously closed tasks automatically.
Upgrading Existing Deployments
Ship the backend and agent/CLI images together. Publish the bundled catalog with
bun run --cwd apps/backend agent-packs:sync, then use its returned revision to preview and apply
an instruction-only pin upgrade for the affected organization:
bun run --cwd apps/backend agent-profiles:upgrade-instructions plan <organization-id> <pins.json> <revision>
bun run --cwd apps/backend agent-profiles:upgrade-instructions apply <organization-id> <pins.json>The preview freezes profile IDs and old instruction pins. Apply is transactional and idempotent; a changed pin rejects the operation. Current model/runtime choices, concurrency, skill selections, MCP grants, enabled/disabled status and memory stay intact. Changed chat pins invalidate persisted harness sessions so the next turn loads the new instructions. Existing active runs keep their immutable manifests. These deployment commands require administrator shell access, not an agent execution token, and do not add HTTP endpoints.
Promotion and prerequisite invalidation now resume unfinished work in its executable phase.
For tasks left in verify by those earlier transitions, preview exact task references:
bun run --cwd apps/backend tasks:repair-continuations plan <organization-id> <repair.json> GRO-3
bun run --cwd apps/backend tasks:repair-continuations apply <organization-id> <repair.json>Review the reason and source run in the report. The repair changes only proven cases, skips active,
legitimate legacy/learning verification and ambiguous histories, and rejects changes since preview.
It preserves ownership, findings, cycles, blockers, history and budgets. Obsolete queued requests are
cancelled in PostgreSQL, making old Redis wakeups harmless. Failed tasks stay failed and report
retryRequired; resume them through ordinary task retry after inspecting its budget and graph.
Never replace the original tasks or broaden ownership to recover a continuation.
For an obsolete platform-repair question, first deploy and verify the actual repair. Then preview the exact open question with a concrete explanation of the repair:
bun run --cwd apps/backend tasks:repair-handoff plan <organization-id> <handoff.json> <inbox-item-id> "Describe the verified platform repair"
bun run --cwd apps/backend tasks:repair-handoff apply <organization-id> <handoff.json>This administrator command cancels the obsolete question without supplying a human answer. It retains an audit receipt and queues the same task, owner and executable phase only if no other blocker remains. It preserves run budgets and dispatch stops. Preview changes to the question, task, profile or authorization configuration invalidate apply. Reapplying the same report is safe, including after a failed Redis wakeup. Genuine business decisions still need the owner’s answer; this command does not reopen terminal tasks or grant integration access. Further repairs use the original coordinator’s bounded plan and existing task/action identities.
After rollout, verify the actual source run/path references from the affected department’s execution capability and compare retained file hashes where receipts exist. Code deployment alone is not proof that an individual historical file is available or that the original task has resumed.
Reading organization knowledge
New-chat suggestions sit below the composer and keep at most four ideas, one per enabled profile.
/organizations/:organizationId/suggestions displays all current verified suggestions with team
filtering, rationale, freshness and supporting-work links. Selection navigates with profile/key/revision
references only. New chat resolves the reference against current same-organization learning and seeds
an editable draft once; it never sends automatically. Replaced/removed/disabled suggestions produce
an unavailable message, and background learning refreshes do not replace typed text.
Context opens with readable facts, preferences and workflows from the published revision. The owner can switch organization or department scope, inspect sources and retained files, search reusable findings, and read the latest verified lessons and suggestions. An unavailable health assessment does not hide readable business Context. The read-only summary endpoint verifies each published document hash and reports document failures separately.
Suggestions are replaced by the next verified learning cycle, including an empty result.
Selecting and inspecting work
Work, notification and approval tables and chat lists expose selection in their rows. The header checkbox selects the visible page; selected IDs remain selected across pagination and reset when the organization or filters change. Bulk actions appear after selection. Selecting all matching filters is explicit, and the existing frozen preview shows the affected items before confirmation. A recent-operation control reopens durable progress. Acknowledged approvals use explicit selection because their view spans multiple statuses.
Work titles truncate in tables with keyboard and hover tooltips. Detail pages show the complete title and a full-width description, with Read more after eight rendered lines. The hierarchy and current-work cards stay 360 pixels tall and scroll internally; each can open independently in a large sheet using the same content.
Instruction Updates On Upgrade
Boot catalog publication reconciles built-in instruction revisions. A canonical profile uses
config.catalogSource.updatePolicy: "auto"; legacy canonical profiles have the same behavior.
Set "pinned" through the profile API to retain a chosen revision. Custom instruction URIs are
never replaced automatically. Publication verifies instruction and selected-skill content before
updating a profile. Incompatible selections retain their old revision and emit an operational issue.
The update preserves profile IDs, status, runtime/model choices, selected skills, grants and memory. Running context manifests retain their original pins. New runs use the new revision; Chat starts a fresh harness session after its profile changes. Explicit administrator preview/apply commands remain available for deliberate instruction migrations.