Threaded WebAssembly doesn't crash. It hangs.
I loaded an Emscripten pthreads build in Chromium, Firefox and WebKit under every isolation-header setup. Seven of twelve hung, and not one rejected. Why each one hangs, how to detect it in milliseconds, and a loader that falls back instead of freezing.
The first time the threaded build of my solver failed, the app sat on "Warming up the solver" forever. The try/catch around the init never fired, and the spinner just kept spinning.
I found the cause and fixed it, but I still didn't know what else could fail the same way. So I built a small test page, loaded the same Emscripten pthreads build in Chromium, Firefox and WebKit, and changed one header at a time. This post is what came out of it: what each failure looks like, why the promise never settles, a check that catches it in a few milliseconds, and numbers on whether the threads were worth it.
The build is tinyfem, the C++ finite element solver behind Elementarium, compiled with Emscripten 6.0.3. The header rules apply to any WebAssembly that uses threads, Rust with wasm-bindgen-rayon included. The exact way it hangs is Emscripten's.
What a threaded build needs
A pthreads build only starts when all four of these hold:
- It's linked with
-pthread(or-fopenmp, which implies it). Its memory becomes aSharedArrayBuffer, so the main thread and the workers can all see the same data. - The page is cross-origin isolated. Browsers only allow
SharedArrayBufferin this mode. You turn it on with two response headers on the HTML page:Cross-Origin-Opener-Policy(COOP) set tosame-origin, andCross-Origin-Embedder-Policy(COEP) set torequire-corporcredentialless. More on choosing between those two below. - The worker scripts are allowed to load. Every thread is a Web Worker, and on an isolated page each worker's script needs its own COEP header.
- The thread pool can start. Emscripten waits for its workers to load the module and report back before your
awaitreturns.
I build two versions from the same C++ sources: a threaded one and a plain single-threaded one to fall back to. The loader later in this post expects both to be ES modules with a default export, which these flags give you:
# single-threaded
emcc ... -O3 -sMODULARIZE -sEXPORT_ES6 -o public/solver/serial/solver.js
# threaded
emcc ... -O3 -sMODULARIZE -sEXPORT_ES6 -pthread \
-sPTHREAD_POOL_SIZE=navigator.hardwareConcurrency \
-o public/solver/threaded/solver.js
Each command writes a solver.js and a solver.wasm. I keep them in the static public/ folder, where the bundler copies them as they are. That keeps the .js, the .wasm and the probe script from later in this post next to each other, so one header rule on /solver/* covers all of them.
The test
I served the threaded build from a small Node server that sets headers per URL, and loaded it with Playwright in Chromium 151, Firefox 153 and WebKit 26.5. The page checks whether it's isolated, then waits up to 10 seconds for the module to load.
| Headers | Chromium | Firefox | WebKit (Safari's engine) |
|---|---|---|---|
| None | hangs | hangs | hangs |
COOP + COEP credentialless, HTML only | hangs | hangs | hangs |
COOP + COEP credentialless, HTML and JS | ready, 169 ms | ready, 265 ms | hangs |
COOP + COEP require-corp, HTML and JS | ready, 154 ms | ready, 239 ms | ready, 154 ms |
"Hangs" means the promise neither resolved nor rejected within 10 seconds. Three different causes are behind those seven hangs.
Hang 1: the page isn't isolated
Without the headers there's no SharedArrayBuffer. Emscripten still creates its workers, one per hardware thread, and then tries to send each one the shared memory. That postMessage throws:
- Chromium:
SharedArrayBuffer transfer requires self.crossOriginIsolated. - Firefox:
The WebAssembly.Memory object cannot be serialized. - WebKit:
DataCloneError: The object can not be cloned.
That error does show up in the console, as an uncaught error. But it's thrown in a callback outside the init promise, so your await never hears about it and your catch never runs. If it gets lost among other console messages, all you have left is a spinner.
So never load the threaded build on a page that can't run it:
if (!globalThis.crossOriginIsolated) {
// load the single-threaded build instead
}
Check crossOriginIsolated in the console on every environment you deploy to. Mine was false on my dev server for months. My security middleware used different COEP defaults in development and production, and production worked, so I never looked. Later I moved the headers into my framework's route config, where they looked right, and the middleware quietly overrode them. Now I check the response headers in the Network tab instead of trusting the config.
Hang 2: the page is isolated, the workers aren't
This one took me the longest, because everything checks out. crossOriginIsolated is true, SharedArrayBuffer exists, and the init still hangs.
On an isolated page, a dedicated worker doesn't inherit COEP from the page. The browser checks the worker script's own response, and if that response has no compatible COEP header, the browser treats it as a network error, even for same-origin scripts. In Chromium you get:
(blocked:response)on the worker script in the Network tab, ornet::ERR_BLOCKED_BY_RESPONSE.- An
errorevent on eachWorkerwith an empty message, no filename and line 0. - One console line from Emscripten:
worker sent an error! undefined:undefined: undefined.
Firefox showed the same 16 empty error events. In both, the promise never settled, because Emscripten keeps waiting for workers that will never start.
You might think this only concerns the worker you wrote yourself. But Emscripten starts its threads from its own generated .js file, the one sitting next to the .wasm:
new Worker(new URL('solver.js', import.meta.url), { type: 'module', name: 'em-pthread' })
Strictly, only scripts that run as workers need COEP. In practice it's simplest to send it on every JS file, so you can't miss one. This is the set that worked in all three engines:
# the HTML page
Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp
# every .js file (and the .wasm, it doesn't hurt)
Cross-Origin-Embedder-Policy: require-corp
Cross-Origin-Resource-Policy: same-origin
Most frameworks and security middleware only add headers to HTML responses, so expect to configure the JS files separately. In my Nuxt app, the HTML headers come from nuxt-security. In development, Vite serves the JS, so it needs its own copy:
// nuxt.config.ts (vite.config.ts works the same way)
vite: {
server: {
headers: {
'Cross-Origin-Embedder-Policy': 'require-corp',
'Cross-Origin-Resource-Policy': 'same-origin',
},
},
},
In production my site is a static build, so the headers live in the host's config instead. Every server that hands out your JS needs checking: dev server, preview server, CDN.
Hang 3: a thread nobody can start
I haven't hit this one myself, but the Emscripten docs are clear about it. A new worker can only start after the thread that asked for it returns to the event loop. If your code creates a thread and then waits for it without returning, for example pthread_create followed by pthread_join, the worker never starts and the wait never ends.
That's what -sPTHREAD_POOL_SIZE in the build above prevents: it starts the workers before your code runs. The value can be a JavaScript expression, and navigator.hardwareConcurrency makes the pool big enough for any thread count you choose later. If your code has to block the browser's main thread, also look at -sPROXY_TO_PTHREAD, which runs your main() in a worker.
Safari and the choice between COEP values
COEP has two values that turn on isolation. They differ in how they treat resources from other origins, like images from a CDN:
require-corpblocks any cross-origin resource that doesn't opt in, either by sending aCross-Origin-Resource-Policyheader or by being loaded with CORS.credentiallessloads those resources anyway, but without cookies.
credentialless is much easier to adopt, because images from a CDN or a storage bucket keep working without changes. That's why I started with it.
But WebKit doesn't support credentialless, on macOS or iOS. It ignores the header, the page isn't isolated, and every Safari user lands in hang 1. With the crossOriginIsolated check in place, they quietly get the single-threaded build instead. Nothing breaks, so you won't notice unless you test in Safari. I only found out by running this test.
So the choice is:
credentialless: Chromium and Firefox get threads, Safari gets the single-threaded build, and your cross-origin images keep working.require-corp: all three get threads, but every cross-origin image, font and script must send CORP, or be loaded withcrossoriginfrom a server that sends CORS headers.
If you choose require-corp, watch out for components that load images themselves. My avatars broke only in the environment that sent require-corp. The UI library preloaded each picture with new Image() and copied only the crossOrigin prop onto it, so the crossorigin="anonymous" attribute I'd put on the component never reached the request. The error says what was blocked, not which code requested it:
ERR_BLOCKED_BY_RESPONSE.NotSameOriginAfterDefaultedToSameOriginByCoep
If your host doesn't let you set headers at all (GitHub Pages, for example), coi-serviceworker adds them from a service worker. It needs one reload on the first visit to take effect.
Detect it in milliseconds
Checking crossOriginIsolated catches hang 1 and the Safari case, but not hang 2, because there the page is isolated. For that, I start a tiny test worker before loading anything big. It sits next to the solver files, so it gets the same headers:
// public/solver/threaded/probe.js
postMessage(self.crossOriginIsolated)
/** Resolves true if a worker can start from `url` and is cross-origin isolated. */
function canStartIsolatedWorker(url: string, timeoutMs = 3000): Promise<boolean> {
return new Promise((resolve) => {
let worker: Worker
try {
worker = new Worker(url, { type: 'module' })
}
catch {
// e.g. a CSP worker-src rule: the constructor throws instead of firing `error`
return resolve(false)
}
const finish = (ok: boolean) => {
clearTimeout(timer)
worker.terminate()
resolve(ok)
}
const timer = setTimeout(() => finish(false), timeoutMs)
worker.onmessage = e => finish(e.data === true)
worker.onerror = () => finish(false)
})
}
On the same twelve setups, the probe gave the same answer as the full 10-second test every time, and it answered in 4 to 106 ms. A blocked worker fires error almost immediately, so you find out before downloading megabytes of wasm.
The loader
Putting it all together, the loader checks isolation, probes a worker, and only then loads the threaded build. That last step still gets a deadline, because a browser might refuse workers for a reason I haven't tested. Anything that goes wrong falls back to the single-threaded build:
const THREADED_INIT_DEADLINE_MS = 8000
/** Reject if `p` hasn't settled within `ms`. Never leaves the timer running. */
async function withDeadline<T>(p: Promise<T>, ms: number, label: string): Promise<T> {
let timer: ReturnType<typeof setTimeout> | undefined
try {
return await Promise.race([
p,
new Promise<never>((_, reject) => {
timer = setTimeout(() => reject(new Error(`${label} did not settle within ${ms} ms`)), ms)
}),
])
}
finally {
clearTimeout(timer)
}
}
// files in public/ are served as-is, so tell the bundler not to touch these imports
const load = (path: string) => import(/* @vite-ignore */ path).then(m => m.default())
export async function loadSolver(): Promise<{ module: any, threaded: boolean }> {
try {
if (!globalThis.crossOriginIsolated)
throw new Error('page is not cross-origin isolated. Check COOP/COEP on the HTML (Safari needs COEP: require-corp)')
if (!(await canStartIsolatedWorker('/solver/threaded/probe.js')))
throw new Error('worker scripts are blocked. Look for (blocked:response) in the Network tab: a .js file is missing its COEP header')
const module = await withDeadline(load('/solver/threaded/solver.js'), THREADED_INIT_DEADLINE_MS, 'threaded init')
return { module, threaded: true }
}
catch (err) {
console.warn('[solver] Using the single-threaded build:', err instanceof Error ? err.message : err)
}
return { module: await load('/solver/serial/solver.js'), threaded: false }
}
The warnings are long on purpose. Nothing in my UI shows whether a solve ran on one thread or eight, so without the warning a fallback just looks like a slow app. With it, whoever opens the console next is told exactly where to look.
With the two checks in front, the deadline should never fire. The build is normally ready in 150 to 270 ms, so 8 seconds leaves room for slow devices. If it does fire, the abandoned attempt may leave idle workers behind. I accept that, because it should be rare and a reload clears it.
Because the two builds are separate files behind dynamic import(), users on the fallback never download the threaded build. If you have several workers, send them all through the same loader so they pick the same build, and the second one comes from the HTTP cache.
Was it worth it?
Here's the speedup on a real workload: a shell plate with about 48,000 unknowns, solved with conjugate gradients inside a Web Worker, which is how my app runs it. Each point is the best of five runs, alternating between the two builds:
Show the numbers
| threads | serial | threaded | speedup |
|---|---|---|---|
| 1 | 1976 ms | 2035 ms | 0.97× |
| 2 | 1967 ms | 1237 ms | 1.59× |
| 3 | 2016 ms | 1005 ms | 2.01× |
| 4 | 2045 ms | 907 ms | 2.25× |
| 5 | 2093 ms | 861 ms | 2.43× |
| 6 | 2090 ms | 858 ms | 2.44× |
| 7 | 2111 ms | 865 ms | 2.44× |
| 8 | 2175 ms | 934 ms | 2.33× |
| 9 | 2387 ms | 999 ms | 2.39× |
| 10 | 2369 ms | 1042 ms | 2.27× |
| 11 | 2400 ms | 1017 ms | 2.36× |
| 12 | 2413 ms | 1051 ms | 2.30× |
| 13 | 2407 ms | 1024 ms | 2.35× |
| 14 | 2369 ms | 1109 ms | 2.14× |
| 15 | 2392 ms | 1261 ms | 1.90× |
| 16 | 2465 ms | 1719 ms | 1.43× |
Yes, but less than the core count suggests. Two threads give 1.6×, four give 2.25×, and then the curve flattens. Anything from 5 to 13 threads sits between 2.3× and 2.44×. After that it falls, and with all 16 hardware threads it's down to 1.43×.
My reading of why, which I haven't proven: most of the solve is conjugate gradient iterations, and each one is mostly a sparse matrix-vector product. That reads the whole matrix from memory and does very little arithmetic per number, so after a few threads they're all waiting on the same memory bandwidth. Past 8, the extra threads don't bring new cores, they share the existing ones. At 16 they also compete with the browser's own threads.
A few practical things came out of these runs:
- Don't trust the default thread count. In all three engines, OpenMP started with 4 threads, whatever the machine had. Here that gets you 2.25× out of a possible 2.44×, which is fine. On a different machine or workload it may not be.
- Don't use
navigator.hardwareConcurrencyas the thread count. It counts hardware threads, not cores. On this laptop it says 16, and 16 is one of the slowest settings. Half of it is a reasonable starting point on machines with two threads per core. Then measure, and let users change it. - Keep heavy threaded work off the page's main thread. I first ran the same benchmark on the main thread, and at 16 threads it was 0.58×, slower than not using threads at all. In a worker, the same setting gave 1.43×.
- Use the single-threaded build when you only want one thread. On one thread the threaded build was 3% slower than the plain one.
- Check that the answers match. A race condition in parallel code can give a slightly different result instead of crashing. I compare the threaded result with the single-threaded one in every run, and here they're bit-for-bit identical. They won't always be, because adding numbers in a different order can change the last digits. If they're not identical, decide up front how close counts as correct.
Checklist
- Build a single-threaded version too, and keep both builds' files together in a static folder.
- Send COOP and COEP on the HTML, and COEP and CORP on every JS file. Confirm in the Network tab, on every environment.
- Use
require-corpif Safari users should get threads. Usecredentiallessif your cross-origin images make that too hard, and accept the single-threaded build on Safari. - Check
crossOriginIsolated, then probe a worker, before loading the threaded build. - Pre-size the thread pool with
PTHREAD_POOL_SIZE. - Never await the threaded init without a deadline, and warn loudly when you fall back.
- Choose the thread count yourself, and measure it on your own workload.
If you've hit a failure that isn't here, I'd like to hear about it: jan@vorisek.me. I'll add it to the post.