Terminal Emulator

Terminal Emulator

active Updated 2026-08-17
javascript dom-security xss-hardening

The site's interactive command-line overlay — and a real DOM XSS I introduced and then fixed, documented here as the security case study it actually is.

Overview

Every page on this site ships a small interactive terminal (` or Ctrl+` to open it) — about 20 commands (ls, cd, whoami, neofetch, search, plus a handful of easter eggs like cowsay and fortune) that read from the site’s own generated JSON data and render colorized output using <span class="term-*"> markup. It’s a toy, but it’s a toy that writes untrusted-shaped strings into the DOM on every keystroke, which is exactly the kind of surface that grows a real bug.

Problem

The output renderer originally did this:

line.innerHTML = text;

text is the return value of whatever command just ran. Most commands only interpolate site data the owner controls (post titles, project names), but a few — cowsay <message>, search <term>, the “unknown command” echo — interpolate the visitor’s own typed input straight into that HTML string with no escaping. Type cowsay <img src=x onerror=alert(1)> and it runs, because innerHTML doesn’t care where the markup came from.

Role

Spotted the direct innerHTML assignment and scoped the fix as PR #24; the actual patch was generated by GitHub’s Copilot coding agent and merged after review. Worth saying plainly rather than implying otherwise — the interesting part of this case study is the sanitizer design, not who typed it.

Approach

Fixing this properly meant not touching the ~20 command handlers, since every one of them already builds its return value as an HTML string sprinkled with legitimate <span class="term-accent">/<span class="term-error"> formatting — auditing each one for escaping bugs would be tedious and easy to regress. Instead, the fix moved to the single chokepoint where every command’s output actually reaches the DOM: appendOutput().

sanitizeTerminalHtml() HTML-escapes the entire string first (so <img src=x onerror=...> becomes inert text), then selectively re-opens exactly two things: <br>, and <span class="term-X"> where X is validated against /^term-[a-z0-9-]+$/ — anything else collapses to a bare <span>. </span> always passes through.

function sanitizeTerminalHtml(html) {
  const escaped = escapeHtml(String(html ?? ''));
  const withSafeBreaks = escaped.replace(/&lt;br\s*\/?&gt;/gi, '<br>');
  const withSafeOpenSpans = withSafeBreaks.replace(
    /&lt;span class=(?:&quot;|&#39;)([^"'<>]+)(?:&quot;|&#39;)&gt;/gi,
    (_, classes) => {
      const safeClasses = classes.split(/\s+/).filter(c => /^term-[a-z0-9-]+$/i.test(c));
      return safeClasses.length ? `<span class="${safeClasses.join(' ')}">` : '<span>';
    }
  );
  return withSafeOpenSpans.replace(/&lt;\/span&gt;/gi, '</span>');
}

Commands typed by the user (echoed back before their output) go through a separate, stricter path — escapeHtml() with no re-opening at all — since there’s no legitimate reason for a literal command line to contain formatting markup.

Result

The vulnerable line.innerHTML = text became line.innerHTML = sanitizeTerminalHtml(text). The whole fix is an 18-line addition to one file — no command handler changed, no existing formatting broke, and every future command gets the same protection automatically since it goes through the same render path.

Decisions

Escape-then-allowlist instead of a blocklist (stripping <script>, onerror=, etc.) — a blocklist has to anticipate every dangerous pattern and reliably loses that race; escaping everything and re-opening a narrow, regex-validated allowlist means anything not explicitly permitted is inert by construction, regardless of what it is.

Fixed the render boundary instead of the ~20 command handlers — the handlers keep authoring <span class="term-*">-formatted strings exactly as before (no call-site changes, no risk of missing one), and the security property lives in one function instead of being an invariant every future command has to remember to uphold.

Lessons