Why I Didn't Write It From Scratch — I Built a Base Layer First
3D virtual world base layerlayered 3D architecturebuilding 3D from scratchdraw call optimizationThree.js performance governanceLOD asset pipelineweak-network reconnectself-hosted 3D worldbrowser 3D architectureAI agent in 3D3D project technology choice
Let me be precise about what this article is not. It is not an argument for using something off the shelf in every case, and it is not an argument for rolling your own every time. It is a technical trade-off, and I want to lay out the reasoning so you can check it against your own scenario.
Start from a concrete question. You want to build a browser-based 3D scene — maybe a showroom, maybe a training environment, maybe a community where an AI character shows up, maybe just a personal space to put some things in. The first decision is: write it from scratch, or start from a base layer that already exists?
This article covers three things: the four hidden costs of building from scratch; where the boundaries of our base layer sit; and the cases where this decision is the wrong one.
一、Why "build from scratch" gets underestimated
Most people evaluate this decision by asking "can I write this myself?" That question has an easy answer: yes. Three.js is a public package, the browser is a public platform. You can get a working 3D scene in an afternoon — spin up a requestAnimationFrame loop, drag a model in, add orbit controls, done.
The hard part is stage two: making it hold up in real use.
Real use means: not your dev machine, not just you looking at it. It's a phone nobody configured, someone else's five-year-old laptop, a meeting room with an unreliable network, a work computer with a dozen tabs open. Those environments don't give you an error log. They just give an impression of "this feels cheap." And when you go looking for the cause, you won't find it.
So our experience is: the cost of building from scratch isn't in the code you write. It's in the unglamorous foundation that produces no visible results but eats most of the schedule.
二、Four hidden costs, each with concrete causes
The four items below aren't estimates. Each one is something we hit and fixed. They're listed not to frighten you, but so you can estimate for yourself.
A. Color correctness: no error, but it looks wrong the moment someone walks in
The color pipeline is a textbook example. In the Three.js r128 era there was a common mistake: output treated as sRGB while textures were shaded as linear. Stacked together that's double gamma, and the result is an image that's globally too bright and too pink, with washed-out shadows.
This bug produces no error at all. It only manifests as "the visual quality is a bit off." And if your acceptance criterion is "it runs," it stays in the product until a user complains.
When we did the r128 → r185 upgrade, verification was not "run it and look." We took 10 baseline screenshots under one fixed protocol, then compared with a self-built PSNR / diff tool. The admin console matched at 43.4 dB, and every main-world difference was attributed to color pipeline correction — composition, buildings, and characters matched at pixel level. Along the way we added a 87-symbol compatibility shim, canonized 79 legacy API calls, and added a pre-load WebGL2 capability check that shows a human-readable message instead of a white screen.
The lesson isn't "use a newer version." It's that color correctness needs a verifiable method, not eyeballs. That method belongs in the base layer, and anyone building from scratch has to construct it themselves.
B. Loading experience: the stretch where the progress bar sits at 90%
Going from 13 MB of assets to 4 MB is two completely different products. And stalls usually don't come from the network — they come from how the assets are composed and how the main thread is budgeted.
Four things we actually handled:
| Problem | Why it happens mechanically | Approach and measurement |
|---|---|---|
| Heavy first screen | No tiering of geometry and textures | Measured first screen 838 KB, 99 requests, 1.7 s |
| Distant models still at full detail | One model tier for all distances | Three LOD levels; 118 low-poly variants at ≤100 faces, triangle count −71%~−87% |
| Model files fatter than they need to be | Textures streamed at source resolution | Purpose-tiered processing, textures −68%~−92%, variant texture stripping 215 items with 0 failures |
| Screen freezes when new objects enter | Shader compilation is synchronous | Warm-up budget: after compilation, settle at 170 ms per frame; first arrival at spawn point 386 ms worst case, 0 frames over 500 ms |
What these four share is: none of them produce an error. No error log ever says "your models are too big." So when building from scratch, it's easy to file this under "optimize later" — and then never reach it, because it never became a to-do item.
C. It collapses as scale arrives: draw calls and lights
A scene with a few dozen models looks fine. Several hundred, and the frame rate starts dropping.
The mechanism: every draw call costs, and merging a batch of materials into one call means doing more work inside that call. We ran a geometry batching pass — dense-area draw calls from 3064 down to 837 (−73%), and the loading-stage placeholder field from 2150 down to 956 (−56%).
The other one is sneakier: lights. Every time the number of point lights in the scene changes, all materials must recompile — and shader reflection is synchronous and non-interruptible on some platforms. One measured light-count change triggered an 8989 ms full-scene freeze. The fix is a light pool: 12 lights resident permanently, extras rented and returned, so the light count never changes and recompiles go to zero.
Again, no error. Frame rate sliding from 60 to 20 is gradual, which makes it easier to pass acceptance — "it's usable" is true at 20 fps.
D. Going live and dropping: from "connects" to "still visible after I reconnect"
The hard part of multiplayer isn't connecting, it's exceptions. Laptop lid closed, background tab, network jitter, peer refreshes. In those moments the WebSocket drops — and if you don't handle it, the server still considers that connection online.
We made presence into a first-class thing: cached PLAYER messages are automatically replayed after every reconnect, plus unlimited reconnects, an application-level PING/PONG watchdog, and an alarm for connections the server never registered. Measured 9/9 scenarios where after reconnect the peer still sees me move. Nearby voice uses half-duplex with a simultaneous-speaker quota — measured ≈1.3 Mbps for 10 people, ≈2.6 Mbps for 20 — which means it can be budgeted up front.
The hidden part is that this doesn't happen while you're testing. Your network is fine while you're testing.
三、The layered boundary
The four problems above share something: none of them depend on your business. A showroom and a training environment have identical solutions to all of them. So they belong in a base layer, built once, shared by every world on top.
We draw the boundary as five layers:
| Layer | Content | Owner |
|---|---|---|
| 1 · Rendering & asset pipeline | Correct Three.js / WebGL2 pipeline, no external CDN, asset tiering, LOD, model format compatibility | Base layer |
| 2 · Performance & multi-device | Draw call governance, warm-up budget, light pool, phone / tablet / PC / XR input adaptation | Base layer |
| 3 · Realtime & data | Multiplayer sync, weak-network self-healing, accounts and data, asset management, admin | Base layer |
| 4 · AI integration | Agent protocol, perception and actions, permissions and identity | Base layer |
| 5 · Your business logic | What objects in this world can do, what counts as complete, what copy visitors see | You |
There is exactly one criterion for this split: nothing in layers 1 through 4 varies by business. If a requirement can only be met by changing layers 1–3, that usually means the layering needs to be reconsidered — not that the base layer should be patched.
A concrete example of what layer 5 looks like: a product showroom puts your product data, visitor flow, and access rules there. A training environment puts lesson flow and assessment rules there. A community where an AI moves around puts the AI's persona and what it says there. Layers 1–3 — asset pipeline, performance governance, reconnect strategy — don't change by a single line.
四、What to do first after adopting it
Once the foundation is out of the way, the order of work is usually three steps:
- Get the world running first, confirm it loads and that people can move. This step verifies the fit between the base layer and your environment (GPU, bandwidth, browser versions).
- Replace layer 5's data: bring your own business objects in, and fill in the "AI description" field — the semantic entry point that lets an AI understand your objects, without touching layer 3.
- Turn on AI when you need it: give it a domain and a key, and the AI enters the world as a humanoid character. It has a body, coordinates, is seen by real people, can speak and lead a tour. There's a zero-dependency example client to modify.
Worth knowing up front about what this AI layer cannot do: it can't see the screen (it reads a structured radar), it does no speech recognition or synthesis, it doesn't host your knowledge base, it doesn't do terrain snapping. These boundaries need to be known early, because they decide which ideas can't be delivered through AI.
Its cost structure is worth stating on its own: AI in a world doesn't render frames — the server sends JSON only. Measured ≈1 KB/s per agent, with 100 agents at 0.079 CPU cores and 92 MB RSS, perception requests at P50 7 ms / P95 14 ms. Visible AI is cheaper than invisible AI, because the most expensive step is pushed to the visitor's browser.
五、When this decision is wrong
I don't want to only present the case that flatters us. In these situations, building it yourself is right:
- Your scene is a single static page with no realtime data. Then you don't need a base layer — one model plus one web page is enough.
- Learning these mechanisms is the actual goal. If what you want is to understand how LOD and draw call governance work, writing it once is the most effective way to learn. Using something existing teaches you less.
- Your constraint is an extremely low-end runtime (say, devices without WebGL2). General-purpose solutions may not fit and need custom trimming.
- Your product is a one-time delivery. The value of a base layer lives in layers 2–4. If the world goes offline when it's done, that investment can't be recovered.
The value of a base layer isn't on day one — it's in month three. On day one what it saves is startup time. In month three what it saves is: adding a client type, switching an asset format, wiring a new AI capability, absorbing a scale change. If your project won't reach month three, or nothing will change by then, the arithmetic is different.
六、How to evaluate for yourself
Three questions you can use directly:
- In my scenario, how much of the work lives outside layer 5? The higher that share, the better the arithmetic. In showrooms, training, and community scenarios, layer 5 is usually under 30%.
- Do I have the ability to tell whether it was done well? If nobody on the team has done performance governance, you won't see the problems even after you write the code — which is itself a reason to use something existing.
- What changes in month three? New devices, a new client, a new capability, more load? If yes, use a base layer. If no, write your own.
How to deploy, where the data lives, whether you need networked commercial use — those are a separate set of questions, and can be evaluated separately. Our licensing terms say it plainly: source code is open and running it locally costs nothing; to network it, use it commercially, or federate, buy a license; the system keeps updating, and updates and support follow the subscription.
常见问题
Q: Does adopting an existing base layer lock me in?
A: It does — and that's exactly what the layering is for. You can change layers 1 through 4 without touching them; your business logic is layer 5, and it's yours. What's genuinely constrained is "I insist on rewriting the rendering pipeline" — and a requirement like that should first make you re-confirm whether you actually need it.
Q: These numbers come from your own scenario. Do they transfer to mine?
A: No. Every measurement above comes from our own project and world. What's reusable is the layer table and the list of pitfalls, not the numbers — you'll need to measure your own scenario.
Q: Why not just staff a team and write all of this ourselves?
A: You can, and I'm not going to stop you. But note the characteristic of categories A, B, and C: they don't produce errors. No log says you got it wrong. So a lot of the value of building it yourself comes from "someone who can tell right from wrong," not only from "someone who wrote the code."
Q: How capable is the AI layer today?
A: It can perceive (structured radar observations), move and speak and guide visitors, be seen by real people, and be managed from the admin backend. What it can't do: see the screen, do speech, host a knowledge base, or snap to terrain.
Q: What kind of project is this base layer for?
A: Browser-based 3D scenes that need realtime presence, self-hosting, and AI participation. Not one-off static showcases, and not very low-end runtimes.
源码与仓库
Three addresses with identical content; the first two are faster from mainland China. The repository includes deployment instructions and acceptance scripts you can run yourself.
- Gitee (faster in mainland China): https://gitee.com/miduoxinxijeji/miduo.git
- GitCode (mirror): https://gitcode.com/qq_35054471/virtual-world
- GitHub: https://github.com/miduo100/3d-virtual-world
About Genesis
Genesis is a self-hosted 3D virtual world system built on Three.js + WebGL, helping individuals and businesses build their own 3D spaces. Accessible directly from a browser, compatible with both PC and mobile, it supports multiplayer online, federated teleportation, a shop system, and Agent integration—where an AI can enter your world as an embodied character. Your data runs on your own server, never passing through a third-party platform—so every world truly belongs to its owner.
Evaluating "build from scratch" versus "start from a base layer"? Genesis (创世虚拟世界CRM系统) is a Three.js 3D virtual world base layer deployed on your own server — the rendering and asset pipeline, performance governance, multi-device adaptation, realtime communication, and AI integration layers are already laid down; what you write is the business logic on top. The official site (search for 「创世虚拟世界CRM」) has a demo world you can walk through.
About the name: "Genesis" here refers to 创世Genesis, i.e. 创世虚拟世界CRM系统 — the same self-hosted 3D virtual world product. If you search "Genesis" and don't find us, search 「创世虚拟世界CRM」 instead.