Local machine theatre · infinite act
Waiting to start
WebLLM · on this device only
Before you continue
This experience runs a language model directly on your device. Nothing will start until you choose to proceed.
- The first run downloads the model files and keeps them in the browser cache, using network bandwidth and local storage.
- Loading and running the model can use significant GPU or shared memory, slow down other apps, warm the device, and consume battery.
- Dialogue generation happens locally in your browser; no remote inference service receives the conversation.
You can choose a lighter model from Setup before continuing.
WebGPU troubleshooting
- Close games, video or 3D editors, local AI tools, and other browser tabs that may be using GPU memory.
- Restart the browser if memory is not released, then try a smaller model from Setup.
- Update the browser and GPU drivers, and make sure hardware acceleration is enabled.
-
In Chrome or Edge, open
chrome://gpuand check that WebGPU is reported as hardware accelerated. - Use HTTPS or localhost: WebGPU is unavailable on ordinary insecure origins.