Files
AICC-Protocol/core.md
T
Emil Shanaty e75ba873d8 aicc/0.1: initial draft
- core.md: human-readable specification covering message format,
  session manifest, tool taxonomy (sensors/actuators/generators),
  tool invocation, events, capabilities, time/causality, memory,
  errors, extension rules, security, conformance.
- schemas/aicc.schema.json: JSON Schema for all message types,
  validated with Draft 2020-12.
- LICENSE: MIT.
- README.md: repository layout, design principles, versioning,
  conformance criteria.
2026-08-08 02:40:27 +03:00

18 KiB

AICC Protocol — Core Specification (Draft 0.1)

Overview

AICC (AI-Controlled Character) is a transport-agnostic protocol for connecting language models to interactive virtual environments. It defines how an agent perceives a world, reasons about it, and acts within it through a structured set of sensor, actuator, and generator tools.

The protocol makes no assumptions about:

  • The host engine (Godot, Unity, Unreal, custom simulators, robotics stacks)
  • The reasoning model (any LLM with tool-use capability)
  • The transport layer (HTTP, WebSocket, stdio, gRPC, in-process)
  • The visual style (2D, 3D, text-based, abstract)

It only specifies message formats, tool taxonomy, capability negotiation, and causality semantics.

Design Goals

  1. Engine-agnostic. The same protocol speaks to a game character, a robot, or a sandboxed LLM playroom.
  2. Model-agnostic. Any LLM that supports tool-use / function-calling can act as the reasoning layer. No prompt engineering baked into the wire format.
  3. Transport-agnostic. Wire format is JSON. Transport is pluggable.
  4. Capability-aware. An environment advertises what an agent may do. The agent never has to guess.
  5. Deterministic when needed. Every tool call returns a structured outcome with timestamps, source, and confidence. Replay is a first-class concept.
  6. Extensible. New sensors, actuators, and generators can be added without breaking older agents, provided they follow the extension rules in §11.
  7. Single source of truth. All world data flows through sensor tools. The manifest carries session metadata only — never world state. This makes the contract honest and uniform across any environment.

1. Model of the World

An AICC interaction involves three actors:

  • Agent — the reasoning layer (typically an LLM). Holds goals, plans, and short-term working memory.
  • Environment — the host system that owns world state. Renders, ticks physics, runs NPCs, persists state.
  • Bridge — the thin adapter that translates between the protocol and the environment's native APIs.
+----------------+        +-------------+        +------------------+
|     Agent      |  <-->  |   Bridge    |  <-->  |   Environment    |
| (LLM + tools)  | AICC   | (adapter)   | native| (engine/world)   |
+----------------+        +-------------+        +------------------+

The bridge is implementation-specific and not part of this spec. The agent sees only the AICC contract.

Tick

A tick is one discrete update of the environment. The cadence is defined by the environment (e.g. 60 Hz physics, 10 Hz AI updates, event-driven). Agents do not assume a fixed tick rate — they negotiate it via the session manifest (§3).

Session

A session is the lifetime of one agent connected to one environment. Sessions are identified by a UUID assigned at handshake. Sessions are stateful: tool calls within a session may accumulate history that influences later calls (e.g. episodic memory).

2. Message Format

All messages are UTF-8 JSON. Every message carries a top-level protocol field with the version string "aicc/0.1".

Common envelope

{
  "protocol": "aicc/0.1",
  "type": "...",
  "session_id": "...",
  "message_id": "..."
}
Field Type Notes
protocol string Must be "aicc/0.1" for this version.
type string One of the message types below.
session_id string UUID assigned by the bridge at handshake.
message_id string UUID for this message. Used for correlation.

Message types

type Direction Purpose
session_init bridge -> agent Open a new session, advertise manifest.
session_resume agent -> bridge Resume a session by session_id.
session_close both Graceful close.
manifest_request agent -> bridge Re-fetch the manifest.
tool_call agent -> bridge Invoke a tool.
tool_result bridge -> agent Return outcome of a tool call.
event bridge -> agent Asynchronous notification from environment.
error bridge -> agent Protocol-level or environment error.
heartbeat both Keep-alive ping with optional payload.

Versioning

The protocol string uses semantic versioning (aicc/MAJOR.MINOR).

  • MAJOR bump: any breaking change to message format or tool schema.
  • MINOR bump: additive changes only. Agents must accept messages from a bridge with a higher MINOR (forward compatibility), ignoring unknown optional fields.

3. Session Manifest

Sent in session_init and on manifest_request. Describes the environment and what the agent may do.

{
  "type": "session_init",
  "session_id": "8f3a...",
  "tick_rate_hz": 10,
  "tick_mode": "fixed",
  "world": {
    "name": "testbed_room_01",
    "kind": "3d"
  },
  "capabilities": {
    "sensors":    ["vision", "depth", "proprioception", "memory_query"],
    "actuators":  ["move", "turn", "look_at", "interact", "say"],
    "generators": []
  },
  "tools": [
    {
      "id": "vision",
      "class": "sensor",
      "description": "Return the current first-person RGB frame as base64 PNG.",
      "input_schema":  { "type": "object", "properties": {}, "additionalProperties": false },
      "output_schema": { "type": "object", "properties": { "png_b64": { "type": "string" }, "width": { "type": "integer" }, "height": { "type": "integer" }, "tick": { "type": "integer" } }, "required": ["png_b64", "width", "height", "tick"] }
    }
  ]
}

Note. World state — including the agent's own position, rotation, inventory, and the world's bounds — is never included in the manifest. It is accessible only through the appropriate sensor tools. The manifest is a contract on the session, not a snapshot of the world.

Tick modes

  • fixed — environment advances in regular intervals. Tools see a coherent snapshot.
  • event — environment advances only on tool calls or external triggers.
  • hybrid — fixed physics tick, but agents may request immediate evaluation via a tick_now tool.

4. Tool Taxonomy

Three top-level classes. Every tool belongs to exactly one.

Class Side effect Returns
sensor none Read-only observation of world.
actuator yes Action performed on the world.
generator yes New world content created.

A tool is identified by its id. IDs are scoped to the session.

Standard sensors (initial set)

ID Returns
vision RGB frame from agent's camera.
depth Depth map aligned to vision.
hear Audio events since last call.
smell Active scent zones near the agent.
touch Surface contact info from last physics tick.
proprioception Position, rotation, velocity, health, inventory.
memory_query Episodic memory hits.
inspect Detailed view of a target entity or location.

Standard actuators (initial set)

ID Effect
move Translate along agent-relative or world axes.
turn Rotate yaw / pitch.
look_at Orient camera toward a target.
interact Use / activate / pick up / talk to an entity.
say Emit speech (or text for dialog-only sessions).
wait Skip ticks without action.

Standard generators (initial set)

ID Effect
place_object Spawn a registered prefab at a location.
modify_terrain Edit terrain heightmap / material.
spawn_entity Create a registered entity at a location.

Generators require the can_modify_world capability and a per-tool rate limit. They are optional in the manifest.

5. Tool Invocation

Request (tool_call)

{
  "type": "tool_call",
  "message_id": "...",
  "call_id": "tc_001",
  "tool": "vision",
  "input": {}
}
Field Type Notes
call_id string Agent-assigned. Echoed in tool_result.
tool string Tool id from the manifest.
input object Must validate against the tool's input schema.

Response (tool_result)

{
  "type": "tool_result",
  "message_id": "...",
  "call_id": "tc_001",
  "ok": true,
  "output": { "...": "..." },
  "meta": {
    "tick": 142,
    "latency_ms": 12,
    "source": "render_thread"
  }
}
Field Type Notes
call_id string Echoed from the request.
ok boolean False on tool-level failure.
output object Conforms to the tool's output schema.
error object Present only when ok is false.
meta object Optional environment metadata.

Failure shape

{
  "ok": false,
  "error": {
    "code": "tool_unavailable",
    "message": "Generator 'spawn_entity' requires capability 'can_spawn_entities'.",
    "retryable": false
  }
}

Standard error codes are listed in §10.

6. Events

Asynchronous notifications from the environment. The agent may subscribe to a subset at session init.

{
  "type": "event",
  "topic": "collision",
  "payload": {
    "other": "wall_segment_03",
    "normal": { "x": -1, "y": 0, "z": 0 },
    "impulse": 4.2
  },
  "meta": { "tick": 142 }
}

Reserved topics:

  • tick — fired on every environment tick (only if subscribed).
  • collision — physical contact.
  • audio — sound event outside the agent's request cycle.
  • state_change — entity added, removed, or substantially changed.
  • agent_message — incoming inter-agent communication in multi-agent sessions.

7. Capabilities

Capabilities are advertised in the manifest. They are coarse-grained permissions, not per-entity ACLs. Fine-grained permissions are an extension (see §11).

Model class declaration

The bridge declares which model class the agent runs on. This lets environments tune tick rates, sensor cadence, and event batching to match the agent's expected response time.

{
  "agent_model": {
    "class": "edge_small | edge_medium | cloud_medium | cloud_large",
    "expected_first_token_ms": 800,
    "expected_full_response_ms": 2500
  }
}
Class Typical target First-token budget
edge_small E2B / phone-class, on-device 100-400 ms
edge_medium E4B / 12B on consumer GPU 300-800 ms
cloud_medium 12B-26B MoE via hosted inference 500-1500 ms
cloud_large 31B+ via hosted inference, with reasoning 1500-3000 ms

These are advisory. The bridge uses them to inform tick-rate choice and sensor buffering strategy. They do not constrain the agent — the agent may exceed or beat the budget — but environments should design their default cadence so that a conformant agent of the declared class can participate in real-time interaction.

Reserved capabilities

  • can_observe_world — read any sensor.
  • can_move_self — invoke movement actuators.
  • can_modify_world — invoke generators.
  • can_spawn_entities — generators that create NPCs / objects.
  • can_modify_terrain — generators that change terrain.
  • can_communicate — invoke say and agent_message events.
  • can_persist_state — write to long-term storage.

An agent must not call a tool that requires an absent capability. Bridges should reject such calls with tool_unavailable.

8. Time and Causality

Tick clock

Every message that carries a meta.tick field references the same monotonic counter maintained by the environment. The agent uses this to reason about ordering of observations.

Eventual consistency for events

event messages may arrive between tool_result messages. The agent MUST process events in meta.tick order; out-of-order arrival is the bridge's responsibility to prevent or flag.

Heartbeat

For long sessions, the bridge may send heartbeat every N seconds. Agents may respond with heartbeat of their own. This is also the vehicle for pushing manifest updates (e.g. capability changes).

9. Memory and State

Memory is split:

  • World state — owned by the environment, accessed only via sensors. The manifest carries no world state.
  • Agent memory — owned by the agent. Not transmitted over AICC unless via the optional memory_query sensor, which is environment-side episodic memory tagged with session events.

The protocol does not prescribe how an LLM maintains its context window. That's the agent's concern. The protocol only guarantees that any state observable via a sensor is reproducible from the session.

Single source of truth

Every piece of world data has exactly one way to reach the agent: through a sensor tool. The agent's own position, the world's bounds, the contents of nearby containers, the time of day — all of these are observable only by calling the appropriate sensor. The manifest never duplicates or caches this data. This eliminates a class of bugs where handshake data drifts from real-time sensors, and keeps the contract uniform across single-agent and multi-agent sessions.

10. Errors

Standard error codes returned in tool_result.error or top-level error messages:

Code When
protocol_mismatch Wire version incompatible.
session_expired Session id unknown or closed.
tool_unknown Tool id not in manifest.
tool_unavailable Capability missing or rate limit hit.
invalid_input Input fails schema validation.
execution_failed Tool ran but failed at the environment level.
timeout Tool exceeded environment-defined timeout.
internal_error Unspecified bridge failure.

All errors include retryable: boolean to indicate whether the agent may safely retry the same call.

11. Extension Rules

New tools and capabilities MAY be added in a MINOR version. Additions MUST:

  1. Use a new id not previously registered.
  2. Be advertised in the manifest.
  3. Define input and output schemas as JSON Schema.
  4. Use only documented error codes.

Breaking changes (renaming, removing, changing semantics of existing tools, changing required fields) MUST trigger a MAJOR bump.

Agents SHOULD ignore unknown fields in meta and unknown optional fields in output. Bridges SHOULD accept tool calls with extra fields by ignoring them (unless strict_input: true is set on the tool in the manifest).

12. Security and Trust

  • The bridge authenticates the agent (out of band of this spec).
  • The agent authenticates the bridge (out of band of this spec).
  • Tools that may be expensive or destructive MUST advertise rate limits in the manifest under tool.limits, e.g. {"calls_per_minute": 60}.
  • All transport MUST be encrypted when crossing untrusted networks. AICC itself does not mandate TLS — that is a transport concern.

13. Conformance

A bridge is AICC-conformant for version aicc/0.1 if it:

  1. Emits and accepts the message types in §2.
  2. Validates all tool_call inputs against the manifest.
  3. Returns errors using the codes in §10.
  4. Preserves event ordering per §8.
  5. Passes the conformance scenarios in conformance/scenarios/.

An agent is conformant if it:

  1. Never invokes tools outside its advertised capabilities.
  2. Echoes call_id correctly on result correlation.
  3. Handles all standard error codes.
  4. Respects tick ordering for events.

Appendix A: Example — Walking Around a Room

  1. Bridge -> Agent: session_init with vision, depth, proprioception, move, turn, look_at.
  2. Agent -> Bridge: tool_call vision.
  3. Bridge -> Agent: tool_result with PNG.
  4. Agent reasons, decides to walk forward.
  5. Agent -> Bridge: tool_call move { "forward": 2.0 }.
  6. Bridge -> Agent: tool_result with new proprioception snapshot.
  7. Bridge -> Agent: event collision when wall hit.
  8. Agent -> Bridge: tool_call turn { "yaw": -45 } to redirect.
  9. Agent -> Bridge: tool_call move { "forward": 1.5 }.
  10. Loop continues.

Appendix B: Open Questions (to resolve in 0.2)

  • How are multi-agent sessions coordinated? Sub-sessions or shared session?
  • How are assets (prefabs, models, audio) referenced — by id, URL, or inline binary?
  • Should vision support streaming chunks, or is single-frame enough?
  • Standard for episodic memory schema (events vs summaries vs both)?
  • Should inspect allow arbitrary queries, or a fixed set of detail modes (color, bbox, semantic, full)?
  • Resolved in 0.1.1: model class declaration (see §7).

Status: draft. Not yet published. Internal review only.