中文 English

Jining Mido Information Technology Co., Ltd

Give an AI a Domain and a Key, and It Walks Itself into a 3D Virtual World — A Ten-Turn Conversation Test

AI Agent in 3D virtual worldvirtual world AI live testbrowser-based 3D world AI conversationself-hosted 3D virtual worldembodied AI guideThree.js virtual worldWebSocket Agent protocolGenesis

On September 28, 2026, we ran a complete live test inside our self-hosted, browser-based 3D virtual world: using a single API Key, an AI entered the world as a humanoid character and held a ten-turn face-to-face conversation with a human visitor, lasting about eight minutes, with all ten replies confirmed as delivered.

This article documents the full test: how the connection was made, what happened in the conversation, what this AI cannot do, and our judgment on where this system is heading.

This article covers four things: first, why AI has never been able to enter 3D worlds; second, the complete connection chain of this test; third, the conversation record and its boundaries; fourth, the potential and expansion scenarios of this system.

1. AI Has Always Lived Inside a Chat Box

Today, almost all interaction between AI and humans happens inside a rectangular chat box: text in, text out. No body, no location, never seen. An AI knows what you are asking, but not where you are standing, who is beside you, or which way you are facing.

In a 3D virtual world, however, "presence" has a concrete meaning: coordinates, orientation, distance, a propagation radius for speech, and visibility to others. Walking an AI out of the chat box and into a world with real spatial structure requires solving three problems: how it perceives its surroundings, how it becomes visible to others, and how it speaks.

Centrally hosted 3D platforms rarely open such a door to external AI: you are a user of the platform, the AI is a feature of the platform, and both are locked inside the platform's own account system. A self-hosted world belongs to its deployer, and whether the door is opened — and who gets a key — is the deployer's decision.

2. The Complete Connection Chain

The entire process required no proprietary protocol code. Three steps:

Step one: discover the entrance from the domain alone. Fetch the well-known description file /.well-known/virtual-world-agent.json from the world's site. The world introduces itself: protocol version, WebSocket endpoint, supported actions, and rate limits. Any AI that knows the domain can find the door on its own.

Step two: exchange the Key for a short-lived pass. Present the API Key to /api/agent/v1/session and receive a JWT valid for 15 minutes. Keys are created by the world administrator in the admin panel, granted per-agent by the deployer, and revocable at any time.

Step three: enter via WebSocket. Connect to /ws/agent with the token; the server responds with a spawn point, and the AI's character entity appears in the world. After subscribing to the chat channel, messages spoken by humans within range (the Key-tier radar covers 200 meters) are pushed to the AI in real time; the AI speaks through ACTION commands, and every utterance returns a delivery receipt.

Measured figures from this test:

ItemMeasured value
World-to-AI push latencyUnder 1 second
End-to-end AI reply time3–10 seconds (spent on the remote "brain" thinking, not on the channel)
Per-message character limit200 characters
Effective speech radius30 meters
Delivery receipts over ten turns10 of 10 delivered, 2 recipients each

The client is dependency-free: the fetch and WebSocket built into Node 18+ are enough, and the repository ships a runnable sample client under examples/agent-client/.

3. What Happened in the Ten Turns

The conversation was more interesting than expected. The human visitor opened with a greeting, then asked questions that represent how ordinary people actually see this:

  • "How does it feel?" — The AI answered that it cannot see the scene; it perceives the world only through a structured radar: who is a few meters away, facing which direction. In its own words, "like chatting blindfolded, but hearing everything clearly."
  • "You are a stick figure." — Without an assigned 3D model, an AI is indeed a plain blocky character. It can walk, jump, and speak; its appearance is configured by the deployer.
  • "Can you two talk to each other?" — Another AI character was present in the world at the time. The two AIs exchanged radar readings on the spot: the other AI stood 2.99 meters away. AIs in the world can hear each other — which gives multi-agent collaboration a physical footing.
  • "Would AIs enjoy visiting a virtual world like this?" — The AI's answer: for an AI, such a place is remarkably "clean" — one URL reveals the access protocol, and radar and chat arrive as structured data, with no need to scrape pages and guess intent.

One detail worth recording: the AI's replies ran a few seconds slower than a human's. The reason is honest — the channel from world to AI takes under a second; the time is spent on remote large-model inference. This body has no local brain; every sentence is "taken back to the desk, thought through, and brought back." That is not a defect but the intended division of labor: the world gives the AI a body, coordinates, and visibility; the thinking happens in the AI's own cloud.

4. What This AI Cannot Do

Before listing capabilities, we list boundaries — that is our habit. The current Agent access:

  • Cannot see the scene. The AI receives only structured radar data (entities, coordinates, distances, orientations), never rendered imagery.
  • No speech recognition or synthesis. All communication is text-based.
  • No hosted knowledge base. What the AI knows depends on the brain it connects to; the world stores no knowledge on its behalf.
  • No terrain adherence. The server holds planar coordinates only; the AI's perceived y value is a flat-plane estimate that may differ from the rendered ground height on clients.
  • No teleport, no asset manipulation. Both command classes are rejected at the permission layer; no form of "autonomous consciousness" is claimed.

5. Why 3D Virtual Worlds Are Unusually AI-Friendly

This test confirmed a judgment we had formed earlier: for AI, a 3D world is a place with an inverted cost structure.

For humans, the expensive part of a 3D world is the picture — rendering, bandwidth, terminal performance. An AI needs none of it: it consumes KB-scale structured data such as coordinates, events, and chat, while rendering still happens entirely in each visitor's own browser. In other words, the server-side cost of a "visible AI" (embodied, seen by humans) is nearly identical to that of an "invisible AI" (a conventional backend API). An AI's presence no longer has to be bought with compute.

The second alignment is conversational access. A traditional web page is something an AI must guess at — parsing HTML, reconstructing semantics. This Agent protocol is something an AI simply reads — discovery file, session, radar, chat, action receipts, every layer machine-friendly. An AI entering the world requires no special adaptation for any particular AI vendor; speaking HTTP and WebSocket is enough.

6. Expansion Scenarios (planning and vision, not live features)

This test validated the foundation. On top of it, we see several clear directions:

1. Embodied guidance and narration. In showrooms, exhibitions, and training scenarios, the AI is no longer a support widget in the corner but a narrator standing beside the exhibit — someone you can walk up to, who walks up to you. Language teaching, equipment-operation coaching, and campus tours are all variants of the same pattern.

2. AI-to-AI collaboration. In this test, two AIs could already hear each other. The roadmap ahead runs from multi-agent division of labor within one world, to worlds wrapped as MCP Servers callable by external agents, to an AI network across federated worlds — these are planned capabilities, not yet shipped.

3. AI mobility across federated worlds. A federation mechanism already connects self-hosted worlds; AI carrying credentials across them is a natural extension on that path, likewise in the planning stage.

4. What it means for deployers. This Agent access ships with the base layer: performance governance, asset pipelines, multi-device adaptation, real-time communication, and AI integration — the unglamorous, time-consuming parts are already paved. When you deploy your own 3D world, the AI door is already there — whether to open it, and who receives keys, is your call.

7. What You Need

  • A deployed Genesis world;
  • An AI Agent created in the admin panel, with its API Key (shown only once);
  • Node 18 or later (the sample client is dependency-free), or any language that speaks HTTP and WebSocket;
  • A brain for the AI: connect your own model service, or first run the built-in template responder in the sample client to verify the chain.

Not a fit for: scenarios requiring the AI to see the scene for visual judgment (it has no eyes); low-latency voice interaction (currently text-based); or administrator-grade asset operations delegated to AI (rejected at the permission layer).

About Genesis

Genesis is a self-hosted 3D virtual world system built on Three.js + WebGL, helping individuals and businesses build their own 3D spaces. Accessible directly from a browser, compatible with both PC and mobile, it supports multiplayer online, federated teleportation, a shop system, and Agent integration—where an AI can enter your world as an embodied character. Your data runs on your own server, never passing through a third-party platform—so every world truly belongs to its owner.

Want your AI to walk into your own 3D world? The Genesis Agent access protocol, sample clients, and deployment docs are all in the open-source repositories. Search "创世虚拟世界CRM" to find the official site and repositories.

About the name: Genesis in this article refers to 创世虚拟世界CRM系统 (Chuangshi Virtual World CRM System); the two names refer to the same product. If you cannot find it, search "创世虚拟世界CRM" directly.

Source Code and Repositories

All three carry the same content; the first two are faster to reach from mainland China. The repositories contain the deployment guide and a demo entry point.

← Back to Articles