中文 English

Jining Mido Information Technology Co., Ltd

The Cost Curve of a 3D World Runs Backwards: First, See Which Line the Money Is On

3D world cost structurebandwidth cost optimizationfirst screen payload3D bandwidth budgetserver rendering costAI integration costself-hosted 3D costdraw call optimizationrealtime communication bandwidthperformance engineering

When people evaluate a 3D world project, the first question is usually "how much will this cost?" That question is hard to answer, because the cost structure of a 3D project isn't the linear "more users, more revenue" shape of a typical internet service. It's a mix: one part grows linearly with visitors, one part is roughly constant, and one part gets cheaper as scale grows. This article quotes no prices — it splits cost into those three categories so you can see where your project's money actually goes, and which spending is wasted. The cost structure described here corresponds to the shape of the Genesis foundation: a browser-based 3D world deployed on your own server.

One: The First Category — Cost That Grows Linearly with Visitors

This one is easy to grasp and easy to estimate: every visitor costs bandwidth, compute, and a realtime connection.

Bandwidth is the biggest line, and its composition is unlike other services. The first screen of a 3D world isn't an HTML file plus a few images — it's geometry, textures, model files. A well-optimized world keeps both request count and size down: our demo world's first screen measures roughly 838 KB across 99 requests, loading in 1.7 seconds. An unoptimized one is still pulling resources after ten seconds. Same code, same server — the entire difference lives in the asset pipeline: LOD tiers cut triangles for distant models, textures are compressed by purpose (measured −68%~−92%), and variant textures are stripped (measured 215 items, 0 failures). These decisions directly drive your bandwidth bill.

Compute is mostly eaten by realtime communication and render instructions. Note the division of labor: frames are rendered in the visitor's browser; the server sends instructions and relays state. So what one server can carry depends on "how much does one visitor cost on the server side," not on how strong the visitor's GPU is. A measured reference: once concurrent speakers in proximity voice are capped by a quota, 10 people ≈ 1.3 Mbps and 20 people ≈ 2.6 Mbps — a budgetable number.

The key to this category isn't "can it be saved" — it's "can it be seen." A cost curve without per-item counters can only be guessed at. That's why we turned render-cost attribution into diagnostic interfaces: per-object counting (onBeforeRender), turn-lag attribution (diagTurnLag), and full-scene sweeps (diagTurnSweep). A cost curve you can't see is a curve you cut by instinct.

Two: The Second Category — Nearly Constant Cost, and Most People's Blind Spot

The second category is one-off or staged investment. It doesn't grow with visitors, but it's often the larger half of a project — and the easiest to underestimate.

Engineering effort. Building a browser-based 3D virtual world from scratch, the time sink isn't rendering. It's the unsexy parts: performance work, asset pipelines, multi-device adaptation, realtime communication, AI integration. They share a property: once done, they don't cost again — but getting them done takes a long time. And because this is "development hours" rather than a "server bill," it usually doesn't appear in quotes, so it gets treated as something that should be quick.

There's a stretch here you have to walk yourself. Three examples from our own work, none of which a design document would have surfaced:

  • Adding one point light recompiles the whole scene. When the light count changes, every material recompiles. On the ANGLE/D3D11 path under Windows, shader reflection is synchronous and uninterruptible, at 160~340 ms per program — one light-count change measured an 8989 ms full-scene freeze. The fix is a light pool: 12 permanent lights, rented and returned, so the count never changes and recompilation drops to zero.
  • A new mesh entering the frame stalls the main thread. Shader compilation freezes the picture. The fix is a warm-up budget: un-warmed meshes stay out of the frame, and after compilation getUniforms() runs at a share of 170 ms per frame, spreading the cost. Measured: first entry at the spawn point worst-cases at 386 ms, with zero frames above 500 ms.
  • A white screen during loading is a UX cost. The fix is a full-scene placeholder field: one draw call covering 1,000+ objects (measured 1,071/1,071 coverage), fading and shrinking out over 250 ms once models land.

All three share a trait: they don't raise errors; they just make users feel the work was sloppy. Which is exactly why they're the easiest to miss and the easiest to wave through acceptance — there's no error log for them.

Maintenance is the most commonly forgotten piece of constant cost. Three.js keeps moving, browser policies shift, AI interfaces change. The real cost of a choice isn't today's code — it's "who keeps this running three years from now." That's the difference between treating engineering as a one-time delivery versus a continuing service, and the two framings produce very different numbers.

Three: The Third Category — Cost That Shrinks with Scale

This is the most counterintuitive part, and the one most 3D project pitches skip.

The cost curve runs backwards: a visible AI is cheaper than an invisible one.

In a typical AI service, cost sits on the inference side — a user sends a sentence, the server runs inference, you pay compute. AI access in a 3D world works differently: give it a domain and a key, and the AI enters your world as a humanoid character with a body, coordinates, visible to real people, able to speak and guide — but it does not render frames. The server sends structured JSON (a "radar": what's nearby, where, how far), and frames are rendered by the visitor's browser on its behalf.

So the cost structure inverts:

ItemTypical AI serviceAI character in a world
Frame renderingCarried by server or dedicated GPUCarried by the visitor's browser
Bandwidth per agentTied to context lengthMeasured ≈ 1 KB/s
Server footprint at 100 agentsRises with inference queueingMeasured 0.079 CPU cores, 92 MB RSS
Scaling meansMore inference computeMore visitor bandwidth

At the same 100 AIs, the server-side footprint of putting them inside a 3D world is small enough to be negligible — because the most expensive step, rendering frames, is pushed onto visitors who are already browsing the world. That's the real cost meaning of "AI enters the world," and why it differs in kind from bolting on an AI chat window.

There's a related property: AI needs no vision here. No screenshots, no screen reading — only structured data. Measured observe radar latency is P50 7 ms / P95 14 ms, with zero 429s inside the window. In a 60-second window with 100 agents, 3 keys, and 2 real players present, the human clients still held 60 FPS with 0 console errors.

Four: Four Spending Moves That Look Like Savings

With the cost structure clear, several moves that appear to save money actually push cost somewhere harder to manage:

Move frame rendering to the server. Once frames render server-side, the cost structure snaps back to "more users, faster burn," with no lever left to pull — a GPU is a hard cost.

Back the loading path with an external CDN. The first screen is indeed faster, at the price of intranet and offline environments breaking, versions drifting, and an uncontrollable long-term bill. After we repointed the importmap locally and localized or stubbed six loaders, offline and intranet both work.

Cut the scene to save frame rate. Deleting models when frames drop moves the cost from a technical problem to a product problem. Measured: dense-zone geometry batching took draw calls 3,064 → 837 (−73%), and the placeholder-field scene went 2150 → 956 draw calls (−56%). Performance has a remediation path; deleting content isn't remediation.

Leave the asset directory untidy. We did an orphan-asset cleanup once: 159 orphans and 2.9 GB deleted, taking the directory from 3.1 GB to 390 MB. This kind of thing doesn't affect functionality, but it slows every cross-world first screen and becomes a recurring problem at package and deploy time.

Five: How to Evaluate Your Own Project

Here are four questions you can use directly — more useful than asking for a price quote:

  1. In category one, how many KB and how many requests does your first screen pull? That's the direct input to the bandwidth bill and to the visitor's first impression.
  2. In category two, how many hours are you budgeting for performance, assets, multi-device, and communication respectively? Once you can state numbers, you know what to write yourself versus what to take ready-made. The same list doubles as an acceptance checklist.
  3. Does your cost include anything invisible that keeps growing? Fonts, model variants, logs, chat records, archives — ungoverned, all of these keep growing.
  4. Within constant cost, what's one-off and what's ongoing? Separating them tells you what investment cycle to plan for.

For licensing and deployment costs, that depends on whether you need networked commercial use and federation, and is worth evaluating separately — this article quotes no numbers and pushes no decision either way.

Six: The Limits of These Judgments

  • All measured values here come from our own project and scenes; your asset mix will produce different numbers. What's reusable is the cost classification method and the four questions, not these values.
  • Splitting cost into three categories isn't the only valid scheme — different teams have different structures. But "which parts scale, which don't" must be answered somehow, and a project that can't answer it isn't ready for technical selection.
  • The most valuable thing about constant cost isn't the hours figure itself — it's that "once done, it doesn't cost again." If something costs a little on every visit, it belongs to category one and shouldn't be counted as finished.

FAQ

Q: Is self-hosting always cheaper?

A: Not necessarily. In category one (bandwidth, realtime compute) self-hosting doesn't remove the cost, it just means you pay it; the constant engineering effort has to be absorbed up front. The real dividing line is category three and "can you change the underlying thing whenever you need to" — only when the core is yours do the later optimizations exist.

Q: Why is a visible AI cheaper?

A: Because rendering is pushed to the visitor's browser and the server sends only structured JSON. Measured at ≈1 KB/s per agent, 100 agents use 0.079 CPU cores and 92 MB of memory — the most expensive part, inference, never lands on your server.

Q: Is a performance pass actually worth it?

A: Depends on what it fixes. A light pool removed an 8989 ms full-scene freeze caused by light-count changes; the warm-up budget pushed long frames under 500 ms. These problems persist forever if unaddressed, and no error log will ever point them out.

Q: What does an AI need to enter a world?

A: A domain, a key, and a server that runs. There's a zero-dependency example client. What you do need to think through in advance is the AI's permission boundaries inside your world and the things it can't do — that's a separate design question.

Source & Repository

All three addresses host identical content; the first two are faster within mainland China. The repository includes deployment docs and the acceptance scripts.

  • Gitee (faster in mainland China): https://gitee.com/miduoxinxijeji/miduo.git
  • GitCode (mirror): https://gitcode.com/qq_35054471/virtual-world
  • GitHub: https://github.com/miduo100/3d-virtual-world

About Genesis

Genesis is a self-hosted 3D virtual world system built on Three.js + WebGL, helping individuals and businesses build their own 3D spaces. Accessible directly from a browser, compatible with both PC and mobile, it supports multiplayer online, federated teleportation, a shop system, and Agent integration—where an AI can enter your world as an embodied character. Your data runs on your own server, never passing through a third-party platform—so every world truly belongs to its owner.

Want to see the cost structure clearly before deciding how to build your 3D world? Genesis Virtual World CRM is a Three.js 3D virtual world foundation you deploy on your own server — asset pipelines, performance work, multi-device adaptation, realtime communication, and AI integration are already laid down, so you only write the top layer. The official site (search "Genesis Virtual World CRM") has a demo world you can walk through.

About the name: Genesis in this article refers to Genesis Virtual World CRM — they are the same self-hosted 3D virtual world product. If "Genesis" doesn't turn us up in search, search "Genesis Virtual World CRM" instead.
← Back to Articles