中文 English

Jining Mido Information Technology Co., Ltd

One 3D Frontend for Phone, Tablet, and XR: Tier by Capability, Not by Device

3D multi-device adaptationbrowser 3D mobile performancephone 3D stutterdraw call optimizationLOD asset tieringshader warm-uplight poolWebGL2 compatibilityweak-network reconnectmultiplayer bandwidth budget3D memory leak

When you build a 3D project, "supports phones" usually lands last — it sits in a corner of the requirements doc until after you've shipped the desktop version. Then you start, and discover that rewriting a separate mobile build isn't an option. The desktop-only things (right-click menus, hover tooltips, scroll-wheel zoom, keyboard shortcuts, file drag-and-drop) simply don't exist on touch. And the phone-only problems (memory ceilings, GPU variance, orientation changes, background freeze) never occur on desktop.

This article covers how we handle it: one world, four clients, tiered by capability rather than forked by device. It also states clearly what's finished and what's only "structurally accommodated but not yet measured" — the latter is marked explicitly below rather than glossed over.

一、Why "responsive" is nearly meaningless in 3D

Responsive web design is a layout problem: width changes, flex wraps, font sizes scale, images swap resolution. In 3D, rendering cost isn't determined by layout — it's determined by scene scale.

The same code on a work computer with a dozen tabs and on a four-year-old phone differs by far more than screen size suggests. Not because the screen is small, but because GPU compute differs by an order of magnitude, available memory differs by an order of magnitude, and the WebGL context you can get is more constrained.

So the first thing to settle isn't "how do we adapt to phones" — it's: which things in my project must follow device capability, and which must stay identical?

We draw that line as four tiers, not four devices:

TierWho gets inFrame-budget orientationAsset strategy
HighDiscrete desktop GPU / high-end mobileProtect image qualityFull assets, full quality
MidMainstream desktop, integrated graphics, modern phonesProtect smoothnessLOD active, shadows degraded
LowOld devices, remote desktop, loading phaseProtect entryPlaceholder field holds the scene, models fill in
XRHeadsets, stereo rendering, must hold above 72 HzProtect steady frame rateMost aggressive geometry and effects budget

The point: tiers are computed at runtime, not hardcoded per device type. The same phone may decide differently on office Wi-Fi than in a subway. The common approach combines several signals: GPU renderer string, physical pixel density, a devicePixelRatio ceiling, and a live sample of short frames — start conservative, measure actual frame time, then raise the tier.

二、The pitfalls of each client, and the matching mechanisms

Phone: memory crashes before frame rate does

In 3D on a phone, dropping to 25 fps is survivable. Running out of memory is an outright crash. And leaks don't show up on a dev machine — you refresh a few times and it's clean, the browser's GC wipes up after you.

This class of problem is hidden because it has three shapes: true leaks (objects still referenced, unreachable for GC), unbounded growth (logs, chat history, event buffers appended forever), and incomplete teardown (leaving a scene but textures and geometry still bound on the GPU). The third is most common and most hidden, because it only becomes visible after repeatedly entering and leaving scenes.

The mechanism is full-chain reclamation: video DOM nodes, light pool slots, speech bubbles, and zombie connections are released item by item; textures and geometry are destroyed on scene switch. Worth noting: this class of problem is only findable with tooling, never by eye. In one memory audit we fixed four leaks by comparing occupancy before and after entering a scene.

The second phone-specific problem is touch and gestures. What a mouse does, touch often can't — or the gesture collides with browser defaults (page zoom, scroll, back). So mobile needs an explicit viewport declaration, interception of gestures it shouldn't claim, and a clean separation between "tap" and "drag." That part differs per project; it belongs to your layer 5.

Tablet: the most underestimated tier

Tablets are awkward because they are neither phone nor computer: big screen, viewing distance further away, GPU close to a phone, input that's "touch but can hover" if a stylus is present.

Practically they're handled with the same pipeline as phones. What's worth calling out is that their first-screen problem is more visible — tablets are often opened "casually," so freezing right at entry costs more than on a phone (bigger share of screen, expectations closer to a computer's).

Desktop: don't treat "it runs" as "it's good enough"

The desktop problem isn't compatibility — it's that your dev machine is too good. 60 fps on your machine can be 25 fps on an integrated-graphics laptop, and users won't blame the device; they'll think your work looks cheap.

The cure is to make rendering cost attributable through diagnostics rather than intuition: per-object counting (onBeforeRender), turn-induced stall attribution (diagTurnLag), full-scene sweeps (diagTurnSweep). The practical use of these interfaces is: you know what's stalling before users complain.

XR: the client where a missing frame rate means unusable

For the first three, dropped frames degrade experience. In XR, dropped frames cause discomfort, because stereo rendering demands an order of magnitude more motion continuity.

XR has non-negotiables: frame rate can't drop dynamically (going unstable is worse than being low), stereo rendering costs twice a single eye, and controller input and hand tracking are two entirely different interaction models.

So XR is its own tier in our scheme, not "the top tier of desktop." That's not "we did XR well so it happens to run" — it's a different budget and a different interaction model. Our XR adaptation is structural (the input abstraction and tier framework leave room for it), but to be explicit: we have not yet completed full measured frame-rate results on real XR hardware, so this article gives you no XR frame-rate numbers.

三、What makes one codebase hold up four clients

The section below is the concrete practice behind those pitfalls, all currently running:

Asset tiering is the highest-leverage item. Distant models drop geometry (three LOD levels, 118 low-poly variants at ≤100 faces, triangle count −71%~−87%), textures are processed by purpose (−68%~−92%), and variant textures are stripped (215 items, 0 failures). With all three, the same world runs on weak hardware and on strong hardware.

Draw call governance determines the size ceiling. After geometry batching, dense-area draw calls went from 3064 to 837 (−73%), and the loading-stage placeholder field from 2150 to 956 (−56%). The placeholder field matters disproportionately on weak devices: one draw call covering 1071/1071 objects, with a 250 ms shrink-out as models land — users always see a complete world first, never scattered pieces.

Shader compilation must be budgeted in advance. Un-warmed meshes don't enter the screen; after compiling, they're settled at 170 ms per frame. Measured worst case at spawn-point entry: 386 ms, with 0 frames over 500 ms. On desktop this is an optimization; on weak hardware it decides whether the first glance stutters.

Light compilation cost amplifies on weak hardware. Every change to the number of point lights recompiles all materials. One measured light-count change triggered an 8989 ms full-scene freeze. A light pool (12 resident, rent-and-return) drives that cost to zero.

The loading phase must not be a white screen. A white screen during load is the most direct experience cost — and on weak devices the white-screen period stretches several times longer. We use a full-scene placeholder field with a shrink-out transition.

Bandwidth needs tiering too. Weak devices are often on weak networks at the same time. After asset tiering, the first screen measures 838 KB across 99 requests in 1.7 seconds; untiered worlds are still pulling resources a dozen seconds after entry.

Unsupported capabilities should say something human. When WebGL2 isn't available, show a clear message instead of a white screen or console error spam — and that condition hits far more often on weak hardware than you'd expect.

四、Backgrounding, screen lock, network loss: three death modes specific to multi-client

Phones and tablets get suspended by the system: backgrounded, locked, interrupted by a call. While suspended the WebSocket may be reaped, and on return the connection is gone without you knowing.

  • Reconnecting is not the same as refreshing the page. We cache presence as messages replayed automatically after every reconnect, plus unlimited reconnects and an application-level PING/PONG watchdog. Measured 9/9 scenarios where the peer still sees me move after a reconnect.
  • The server has to know the connection is gone. Alarms and unregistered-connection detection — otherwise your online count is fiction.
  • Resources like voice need quotas and reclamation. Nearby voice uses half-duplex plus a simultaneous-speaker quota, measured ≈1.3 Mbps at 10 people and ≈2.6 Mbps at 20 — directly budgetable.

None of these three tend to happen on desktop, which is exactly why they pass testing — they only appear in genuine multi-client use.

五、Reusing this method for the next client

The value of tiering by capability isn't how many clients we support. It's that you can reuse it for your next client:

What a new client must answerWhere it sits in the tier framework
Where does its GPU capability land?Runtime detection logic — no changes needed
What's its frame-rate floor?Decides whether it's its own tier or a tightened existing one
What's its input model?Input abstraction layer: add a mapping, don't touch rendering
Does it have unique system events? (backgrounding, calls, power-save)Hook into the lifecycle reclamation points

In other words: a client isn't a config file — it's a combination of capability tier + input mapping + lifecycle events. Adding a client means changing those three things, not the whole renderer.

六、Boundaries of this approach

A few things that must be said plainly:

  • We have not completed measured results on mobile and XR. Every number in this article comes from desktop acceptance scripts (the scripts are in the repository; run them yourself). Mobile and XR frame rates, power draw, and thermal behavior need measuring on your own target devices — don't extrapolate our desktop numbers to phones.
  • Tier detection has no universal threshold. How to weight GPU strings, pixel density, and short-frame samples depends on your scenario and audience. What we're giving is a mechanism, not a recipe.
  • Gestures and UI are layer 5. The four tiers above govern render budget and resource reclamation, but whether a tap zooms or a two-finger pinch does depends on your product — that part is always yours to write.
  • Trade-offs on weak hardware are product decisions, not technical ones. Whether your world can even load on an old phone depends on whether you're willing to cut features. Technically it can run; whether the product should is a separate question.
  • We have a known gap of our own: movement is still advanced by a fixed per-frame increment, so on low frame rates the feel is slow. That's written down in our technical notes rather than hidden.

常见问题

Q: Why not just make three separate projects?

A: You can, but the cost is three times the maintenance and three fork risks. The real question is which things must follow device capability — pull those out into tiers, and desktop and mobile share one codebase, with only the thin input-mapping layer actually forking.

Q: How much of the work is multi-client adaptation?

A: I don't have a reliable hour count and won't invent one. What I can give you is a method: walk through the pitfalls in section two against your own scenario and mark which ones you must write yourself. The rest is what the base layer already covers.

Q: What frame rate can phones get?

A: I can't give you a number, because we haven't completed mobile measurements. Desktop measurements are citable (first screen 838 KB / 99 requests / 1.7 s); for phones, run it on your target models. Whether you do asset tiering matters more to phones than which framework you pick.

Q: Does XR take extra work?

A: Yes, and not parameter-level work. The cost of stereo rendering, two distinct input models, and the inability to drop frame rates dynamically make XR closer to its own tier than a high-end desktop. Structurally we've left room for it; measured performance on real hardware is still pending.

Q: Isn't a placeholder field "fake content" — will users read it as cutting corners?

A: It depends on how clean the exit is. The approach is one draw call covering every object (measured 1071/1071) with a 250 ms shrink-out as models land — users see a complete world first, then detail swaps in. Done well it's a plus, because "there's something when I walk in" beats "the world is blank."

源码与仓库

Three addresses with identical content; the first two are faster from mainland China. The repository includes deployment instructions and a batch of acceptance scripts you can run yourself to see the desktop measurements.

  • Gitee (faster in mainland China): https://gitee.com/miduoxinxijeji/miduo.git
  • GitCode (mirror): https://gitcode.com/qq_35054471/virtual-world
  • GitHub: https://github.com/miduo100/3d-virtual-world

About Genesis

Genesis is a self-hosted 3D virtual world system built on Three.js + WebGL, helping individuals and businesses build their own 3D spaces. Accessible directly from a browser, compatible with both PC and mobile, it supports multiplayer online, federated teleportation, a shop system, and Agent integration—where an AI can enter your world as an embodied character. Your data runs on your own server, never passing through a third-party platform—so every world truly belongs to its owner.

Evaluating how to cover multiple clients for a 3D project? Genesis (创世虚拟世界CRM系统) is a Three.js 3D virtual world base layer deployed on your own server — asset tiering, performance governance, multi-client input adaptation, weak-network self-healing, and AI integration are already laid down; what you write is the business logic on top. The official site (search for 「创世虚拟世界CRM」) has a demo world you can walk through.

About the name: "Genesis" here refers to 创世Genesis, i.e. 创世虚拟世界CRM系统 — the same self-hosted 3D virtual world product. If you search "Genesis" and don't find us, search 「创世虚拟世界CRM」 instead.
← Back to Articles