Replication report · interactive SVG figures

How those animated sections are built

The “Inside a harness” and “LLM RL · multi-rollout” sections in FineEnvs/multi-harness-rl are not videos, not GIFs and not screenshots. Each is a self-contained HTML fragment that draws its own SVG, frame by frame, from a small state machine. This report is the decompiled pattern, the reasons behind each choice, the failure modes, and a recipe you can hand to a person or to an AI agent.

Contents

  1. The answer in one paragraph
  2. What was examined
  3. The five layers
  4. The mount layer, in full
  5. The file contract
  6. Geometry: one layout pass
  7. The clock and the state
  8. The render
  9. The controlled variant
  10. Live demos
  11. The recipe
  12. Pitfalls
  13. Briefing an AI agent
  14. Files here

01The answer in one paragraph

Each figure is a self-contained HTML fragment: a root <div>, a <style> scoped by that div's class, and one <script> wrapped in an IIFE that puts nothing on window. The article page injects the fragment's markup at build time, which means the browser arrives at a page whose inline script never executed — injection through innerHTML and equivalents does not run scripts. A separate loader script then finds those dead <script> elements and re-executes their text by hand. Once alive, the fragment takes over: it measures its own text, computes a complete layout, then runs a single requestAnimationFrame loop that advances a state machine and rebuilds the entire SVG as a string every frame. There is no animation library, no CSS keyframes, no D3 in these two figures, and no per-element DOM bookkeeping. The animation is a pure function of state.

02What was examined

All of the following is public in the Space repository (huggingface.co/spaces/FineEnvs/multi-harness-rl). The article is an Astro build, so figures live as separate files rather than inside the prose.

PartPathRole
Loaderapp/src/components/HtmlEmbed.astroInjects the fragment, then re-executes its scripts lazily
“Inside a harness”app/src/content/embeds/d3-harness-anatomy.html13.9 KB, the figure in your screenshot
“LLM RL · multi-rollout”app/src/content/embeds/d3-llm-rl-coding.html27.6 KB, the same engine plus Play / Reset / Speed
Their authoring spec.ai/skills/create-html-embed/A skill written for their own agents: SKILL.md + a 25 KB directives.md
Build stepapp/src/utils/extract-embeds.mjsCollects and inlines the embeds/ fragments
Scaleapp/src/content/embeds/*.html30+ figures sharing one contract

Two observations worth keeping. First, the d3- filename prefix is a convention, not a dependency: the data charts in that article use D3, but both figures discussed here are hand-built SVG strings. Second, the project ships AGENTS.md, CLAUDE.md, GEMINI.md and a .ai/skills/ tree — they treat “write a figure” as an agent task with a written contract. Section 12 generalises that.

03The five layers

LayerQuestion it answersRuns when
1. MountHow does inert markup start running?Once, when the figure nears the viewport
2. ContractWhat shape must the file be?Always — it is the format
3. GeometryWhere does everything go, at this width?On start, resize, and theme change
4. Clock + stateWhat is true right now?Every animation frame
5. RenderWhat does that look like?Every animation frame, from state alone

The layering is what makes the figure robust. A resize is a geometry event, not a redraw realignment. A theme switch is a colour re-read, not a rebuild. A background tab is a clock event, not a rendering event. Each concern has exactly one trigger.

04The mount layer, in full

The Astro component loads every embed file as a raw string at build time and injects it with set:html:

// HtmlEmbed.astro (abridged)
const embeds = import.meta.glob("../content/embeds/**/*.html",
  { query: "?raw", import: "default", eager: true });

<figure class="html-embed" data-embed>
  <div set:html={htmlWithId} />      // injects markup AND the dead <script> tag
</figure>

Injected scripts do not run, so a second component script re-executes them. Their version, trimmed to the load-bearing lines:

const mount  = scriptEl.previousElementSibling;      // the injected fragment
const execute = () => {
  mount.querySelectorAll("script").forEach(old => {
    if (old.dataset.executed === "true") return;      // never run twice
    old.dataset.executed = "true";
    if (old.src) { /* copy attributes into a fresh script appended to body */ }
    else { (0, eval)(old.text || ""); }               // run the inline body
  });
  figure.classList.add("html-embed--loaded");
};

if ("IntersectionObserver" in window) {
  const observer = new IntersectionObserver(entries => {
    entries.forEach(entry => { if (entry.isIntersecting) { observer.disconnect(); execute(); } });
  }, { rootMargin: "..." });
  observer.observe(figure);
  setTimeout(execute, 3000);                          // fallback if it never reports
} else { /* run on DOMContentLoaded */ }

The load-bearing part is the data-executed flag, not the execution mechanism. The source Space re-executes the inline body with (0, eval). That works, but it forces every host that serves the page to allow 'unsafe-eval' in its Content-Security-Policy. The version shipped here injects a <script> element with the body as textContent instead: it runs in the same global scope and needs only 'unsafe-inline', so the policy stays free of unsafe-eval. Either mechanism is made safe by the same flag. It is not defensive style — it is required, because an injected fragment can be mounted more than once (lazy-load retries, client-side navigation, a re-render). Without it the figure boots twice, two requestAnimationFrame loops run against the same DOM, and it animates at double speed while fighting itself.

Why lazy at all

Thirty figures each holding a 60 fps loop would burn a phone battery for content the reader may never reach. The loader mounts a figure only when it is roughly one screen away, and the figure itself then observes its own visibility (the next section) so it stops drawing the moment it scrolls out.

05The file contract

<div class="circuit-embed">                 <!-- root; class name == filename -->
  <div class="ce-canvas"></div>              <!-- mount point for the animated layer -->
</div>
<style>  .circuit-embed { ... }  </style>      <!-- every rule scoped by the root class -->
<script>
(() => {
  const C = 'circuit-embed';
  const boot = () => {
    const nodes = [...document.querySelectorAll('.' + C)].filter(e => e.dataset.mounted !== 'true');
    const r = nodes[nodes.length - 1];        // last unmounted instance
    if (!r) return; r.dataset.mounted = 'true';
    ...
  };
  if (document.readyState === 'loading') document.addEventListener('DOMContentLoaded', boot, { once: true });
  else boot();
})();
</script>

Five rules are doing real work here:

The one bug the source still has: SVG id collisions. The harness figure hard-codes <clipPath id="ha-strip"> and refers to url(#ha-strip). Two instances of that figure on one page share the first id, and the second instance clips its bars to the first instance's rectangle. The fix costs one line and this replication uses it: const uid = 'ce' + Math.random().toString(36).slice(2, 8); and then id="${uid}-strip". Any defs child — filters, gradients, clip paths, markers — needs the suffix.

06Geometry: one layout pass

Before anything moves, the figure computes a complete geometry object. The harness figure's layout() derives the container width, decides between the wide and narrow arrangement, measures each row's text, sums the row heights, and records the total viewBox height. Only then does rendering begin.

The interesting part is the text measurement. SVG has no automatic wrapping, and foreignObject is fragile, so they measure with a canvas context:

const mc  = document.createElement('canvas').getContext('2d');
const fam = getComputedStyle(r).fontFamily || 'system-ui, sans-serif';
const measure = (t, size, weight) => { mc.font = `${weight||400} ${size}px ${fam}`; return mc.measureText(t).width; };

const wrap = (t, size, weight, maxW) => {           // greedy word wrap by measured width
  const words = t.split(' '); const lines = []; let cur = '';
  for (const w of words) {
    const test = cur ? cur + ' ' + w : w;
    if (measure(test, size, weight) > maxW && cur) { lines.push(cur); cur = w; } else cur = test;
  }
  if (cur) lines.push(cur);
  return lines;
};

const rows = STAGES.map(s => ({ ...s, dl: wrap(s.d, 11.5, 400, textW), h: 11 + TX + 6 + wrap(s.d, 11.5, 400, textW).length * 15 + 10 }));

Their own comment above that code explains why it exists: an earlier version divided the box height by the row count and drew two text lines into a box tall enough for one. A hand-drawn diagram has no layout engine behind it; if you do not measure, you overflow. The row height is therefore a function of the measured line count, and the total height is a function of the row heights.

Layout is re-run on exactly two events, and never inside the frame loop:

new MutationObserver(() => { layout(); paint(); })
  .observe(document.documentElement, { attributes: true, attributeFilter: ['data-theme'] });

if ('ResizeObserver' in window) {
  let t; new ResizeObserver(() => { clearTimeout(t); t = setTimeout(() => { layout(); paint(); }, 120); }).observe(r);
} else { window.addEventListener('resize', () => { layout(); paint(); }); }

The 120 ms debounce matters: on a phone, a rubber-band scroll or an address-bar collapse fires resize dozens of times per second, and each one would otherwise trigger a full re-measure and repaint.

07The clock and the state

One loop, one clock, and a state machine advanced by elapsed time:

const STAGE_MS = 820;                    // a stage lasts this long, regardless of frame rate
let stage = 0, phase = 0, pulses = [], trips = [], calls = 0, turn = 1, visible = true;

const step = (dt) => {
  phase += dt;
  trips.forEach(t => t.age += dt);                        // age drives fade-in
  pulses.forEach(p => p.t += dt / TRIP_MS);               // progress 0..1 along the wire
  pulses.filter(p => p.t >= 1 && p.out).forEach(() => pulses.push({ t: 0, out: false }));
  pulses = pulses.filter(p => p.t < 1);

  if (phase >= STAGE_MS) {                                // accumulator, not a counter of frames
    phase -= STAGE_MS;
    if (STAGES[stage].k === 'sandbox') calls++;
    stage = (stage + 1) % STAGES.length;
    if (stage === 0) { turn++; pulses.push({ t: 0, out: true }); trips.push({ age: 0 }); }
  }
};

const tick = (ts) => {
  if (!last) last = ts;
  const dt = Math.min(64, ts - last); last = ts;          // clamp: a stalled tab must not fast-forward
  if (visible) { step(dt); paint(); }
  raf = requestAnimationFrame(tick);
};

Five design decisions in fifteen lines, and each one is a bug avoided:

The controlled variant multiplies the same dt, so speed is one scalar rather than scattered durations:

const step = (dt) => {
  S.t += dt;
  S.phase += dt * S.speed;                                 // 0.25x .. 2x, one multiplier
  const drift = (current() / I_FULL) * 0.00042 * S.speed * dt;
  S.pulses.forEach(p => { p.s = (p.s + drift) % 1; });
  ...
};

08The render

Every frame, the figure builds one SVG string and replaces the animated subtree:

const paint = () => { r.innerHTML = svg() + legend() + note; };

That looks wasteful and is not. For a figure with a few dozen nodes:

Accessibility and theme handling ride along:

let s = `<svg viewBox="0 0 ${W} ${H}" role="img" aria-label="A closed series circuit: ...">`;

// colours are read while drawing, so a repaint is all a theme switch needs
const V = (k, f) => (getComputedStyle(r).getPropertyValue(k).trim() || f);
const c = V('--ce-current', 'rgb(78,165,183)');

Colours come from CSS custom properties read at draw time. That single decision means: the figure inherits the article's palette rather than hard-coding it, dark mode needs only a repaint, and a designer can restyle every figure on the site by editing tokens. Text that originates outside the figure is escaped before interpolation (esc() for &, <, >) — an SVG string built from data is still an injection surface.

Finally, the motion preference is a first-class branch, not an afterthought. Both figures render a representative still when the reader has asked for reduced motion:

const reduce = matchMedia('(prefers-reduced-motion: reduce)').matches;
if (reduce) {
  stage = 2; calls = 3; turn = 4;                          // mid-turn, one request in flight
  trips = Array.from({ length: 6 }, () => ({ age: 999 }));  // some history
  pulses = [{ t: .55, out: true }];
  paint();                                                 // no rAF started at all
  return;
}

09The controlled variant

“LLM RL · multi-rollout” is the same engine with a transport: Play/Pause, Reset, and a Speed range from 0.25× to 2×. The controls are plain HTML authored once in the fragment's static markup — not drawn in SVG, and not rebuilt by paint():

<button type="button" data-act="play"><span data-label>Play</span></button>
<button type="button" data-act="reset">...</button>
<label>Speed <input type="range" min="0.25" max="2" step="0.25" value="1" data-act="speed">
  <span data-speed-val>1.00×</span></label>

The rule that follows from this: never rebuild a control in the per-frame paint. If a <input type="range"> is recreated 60 times a second, the element under the reader's finger is destroyed mid-drag, and the slider jumps, loses focus, or stops responding. Keep controls in the static skeleton; let paint() update only text nodes and classes on them. This replication follows the same split, and adds a physical switch (open/closed) and a resistance slider, because those change the state the whole figure derives from.

10Live demos

Both demos below are the real thing, mounted by the same loader described in section 4. First the minimum viable pattern — geometry, clock, state, render, in about fifty lines:

Demo A. One state machine (three stages), one drifting packet, one repaint per frame. No library.

Then the full figure: a series loop with a source, a conductor, a load, a switch and a meter. Every visible quantity is derived from one value, I = V/R — carrier drift speed, meter bar heights, bulb brightness and the halo all read from it. Break the switch and all of them go to zero together, which is the pedagogical point the animation exists to make.

Demo B. The replication target: src/circuit-embed.html, mounted by the same loader. Try: Switch: closed, then Resistance at 100 Ω.

Verifying that a figure is actually animating

Do not judge by eye at 60 fps, and never accept “it looks animated” as evidence. Read the DOM twice and compare, exactly as a test would:

// in the console: is the painted subtree really changing?
const el = document.querySelector('.circuit-embed svg');
const a = el.outerHTML.length;
const snapshot = () => document.querySelector('.circuit-embed svg')
    .querySelector('circle').getAttribute('cx');     // a carrier's x
const t0 = snapshot();
await new Promise(r => setTimeout(r, 400));
[ 'first sample', t0, 'after 400ms', snapshot(),
  'changed', t0 !== snapshot() ].join(' ');

11The recipe

Build order

  1. Write the state object first. List every value that changes and every value derived from it. If you cannot name the state, you cannot animate it.
  2. Draw one static frame by hand as a literal SVG string with the numbers written out. Get it right at one width before adding a clock.
  3. Extract layout(). Replace the literal numbers with values computed from the container width. Add a narrow branch.
  4. Extract render(). Make it a pure function of state, returning a string. Confirm that paint() at a fixed state reproduces your step-2 frame.
  5. Add the clock. requestAnimationFrame, dt clamped, fixed-duration stages, one accumulator.
  6. Add the lifecycle. Visibility gate, resize debounce, theme observer, reduced-motion still.
  7. Add controls last, in static HTML, wired to state. Never inside the loop.
  8. Mount it, watch it, then test it with the two-reads comparison above.

The twelve rules

  1. One file per figure. Root class equals filename. All CSS prefixed by the root class.
  2. Everything inside an IIFE. Nothing on window.
  3. Mount guard: data-mounted, and select the last unmounted instance.
  4. Unique uid suffix on every SVG id (filters, gradients, clip paths, markers).
  5. Measure text; never guess how many lines a caption needs.
  6. Compute a full layout object outside the frame loop.
  7. Advance state by elapsed milliseconds; clamp dt.
  8. Rebuild the SVG as a string per frame; keep controls out of that subtree.
  9. Read colours from CSS custom properties at draw time.
  10. Escape any interpolated text.
  11. role="img" + a real aria-label; a still frame under prefers-reduced-motion.
  12. Pause when off-screen. Stop on colab stop — whatever keeps your loop alive, it must be able to die.

Definition of done

12Pitfalls

SymptomCauseFix
Figure renders, then never movesMarkup was injected, script was never re-executedRe-execute inline scripts after injection, once, with a data-executed flag
Animates at double speed, twitchesFragment mounted twice, two rAF loops on one DOMThe data-mounted guard plus the loader's executed flag
Runs fast after you return to the tabUnclamped dt after a stalldt = Math.min(64, ts - last)
Second copy clips weirdlyDuplicate SVG id from a shared defsPer-instance uid suffix on every id
Caption overlaps the next rowText wrapped by eye, not measuredCanvas measureText + greedy wrap, height from line count
Slider jumps while draggingControl rebuilt inside the per-frame paintControls live in static markup; paint only updates labels
Layout still wrong after a resizeRendered at a new width but never re-measuredRe-run layout() from a debounced ResizeObserver
Colours stale after a theme switchPalette captured once at bootRead the CSS variables at draw time and repaint on data-theme
Whole article restyled by one figureUnprefixed CSS rule in the fragmentPrefix every selector with the root class

13Briefing an AI agent

The Space authors solved this the same way: a written contract in the repository (.ai/skills/create-html-embed/SKILL.md and its 25 KB directives file) plus AGENTS.md/CLAUDE.md/GEMINI.md so any agent that enters the repo inherits the conventions. Generalised, the brief has six fields. Give all six, or the agent will invent the missing ones.

FieldWhat to state
SubjectWhat the figure explains, in one sentence, and the one idea the reader must leave with
StatesThe named states the machine cycles through, in the order they occur
Derived valuesWhat the reader may change (a control) and which visual quantities follow from it
ContractRoot class, filename, no globals, mount guard, unique ids, controls in static HTML
LifecycleVisibility gate, resize re-layout, theme repaint, reduced-motion still, off-screen pause
EvidenceHow the agent must prove it works: two-reads-differ, 360 px and 1100 px renders, reduced-motion still, no window leakage

A brief that satisfies this, ready to paste:

Build ONE self-contained HTML figure file, no dependencies, no build step.

Subject: how a closed series circuit lights a bulb, and why breaking the
loop anywhere stops the current everywhere.
States, in order: the source -> the conductor -> the load -> the return path.
Control: a switch (open/closed) and a resistance slider (10-100 ohm).
Derived: I = V/R drives carrier drift speed, meter bar heights, and bulb
brightness (remove power when the switch is open).

Contract: root div class == filename; all CSS scoped by that class; one IIFE,
nothing on window; data-mounted mount guard; a unique uid suffix on every SVG
id; controls authored in static HTML and never rebuilt by the frame paint.
Lifecycle: gate on IntersectionObserver, re-layout on debounced ResizeObserver,
repaint on a data-theme mutation, render one meaningful still frame when
prefers-reduced-motion is set, no rAF loop in that case.
Evidence: report the DOM attribute comparison that proves motion, the two
widths you rendered, and confirm no new window properties.

Make the agent's acceptance criteria mechanical. “Animate it nicely” cannot be checked. “Read circle.cx twice 400 ms apart and show the values differ” can. The checklist in section 11 is written to be executable rather than aesthetic, which is the difference between a figure that is finished and a figure that merely looks finished.

14Files here

FileWhat it is
index.htmlThe deliverable page: the animated closed-circuit figure, mounted by the loader
circuit-embed.htmlThe raw embed fragment — the file you would drop into a Space, an MDX page, or any host that re-executes scripts
report.htmlThis report

Reuse: copy circuit-embed.html, rename it, and change the root class to match the new filename. The engine — layout, clock, render, lifecycle — stays.