Skip to content

Cameras, Panning, Zooming and Screen to World

There is no camera. There is a transform applied to everything else, and calling it a camera is a convenience.

Once you accept that, the rest falls out: drawing is one matrix, clicking is its inverse, zooming toward a point is solving a one-line equation, and parallax is the same camera with its position scaled down. This Section closes Part 3, and it is where Section 3.2’s inverse earns its place.

Put the camera at (10,4)(10, 4). Nothing about the world changes — what changes is that everything must be drawn (−10,−4)(-10, -4) from where it sits, to bring that spot to the middle of the screen. Zoom in twice and the world is drawn twice as large, not half.

So the matrix you draw with is the inverse of the camera’s own placement:

V=S(zoom)  R(−θ)  T(−position)V = S(\text{zoom}) \; R(-\theta) \; T(-\text{position})

Every term is the opposite of the camera’s. Read right to left, as 3.1’s conventions require: move the world so the camera sits at the origin, un-rotate it, then scale it up.

Writing it out like that avoids inverting a matrix every frame. But it also means there are now two expressions of the same idea, and they can drift — so the build asserts this hand-written form is exactly inverse(cameraPlacement), entry by entry, across a range of positions, zooms and rotations. Each of those three terms could have its sign flipped independently and still produce a picture that looks like a camera.

The camera does not know a screen exists. It converts world to view coordinates — origin at the middle, Y up. Turning those into pixels is a separate matrix:

viewport=T ⁣(w2,h2)S(1,−1)\text{viewport} = T\!\left(\tfrac{w}{2}, \tfrac{h}{2}\right) S(1, -1)

Centre the origin, then flip Y. That is the whole of Section 1.1, in one place, and keeping it separate is what lets one camera drive a phone and a monitor without knowing which it is.

A pannable, zoomable world in three parallax layers
Move your pointer over the scene to read the world position under it.
The code that draws it src/lib/gamedev/demos/2d/camera.scene.ts
/** A pannable, zoomable world in three parallax layers, with the pixel under the pointer read back. */
import { makeCanvas2D, dot as fillDot, label, line } from "../canvas2d.ts";
// From `controls.ts`, not `ui.ts`: the latter imports Three.js and this track must not.
import { addCheckbox, addReadout, addSlider } from "../controls.ts";
import { worldToScreen } from "../../../gamedev2d/camera2d.ts";
import {
  FLAG,
  LAYERS,
  RANGE,
  cameraFor,
  flagOnScreen,
  layerCamera,
  ridge,
  visible,
  worldUnderPixel,
} from "./camera-shared.ts";
import type { MountFn } from "../runner.ts";

const FLAG_COLOUR = "#f0883e";
const CURSOR = "#d2a8ff";
const GRID = "#1c2229";
const TEXT = "#9198a1";

const mount: MountFn = (el) => {
  const { ctx, canvas, width, height, clear } = makeCanvas2D(el, 330);

  const show = addReadout(el);
  const note = addReadout(el);
  const cameraX = addSlider(
    el,
    "camera x",
    RANGE.cameraX.min,
    RANGE.cameraX.max,
    0,
    draw,
    "",
    0.2,
  );
  const cameraY = addSlider(
    el,
    "camera y",
    RANGE.cameraY.min,
    RANGE.cameraY.max,
    1,
    draw,
    "",
    0.2,
  );
  const zoom = addSlider(
    el,
    "zoom, pixels per unit",
    RANGE.zoom.min,
    RANGE.zoom.max,
    RANGE.zoom.min,
    draw,
    " px",
    1,
  );
  const anchorZoom = addCheckbox(
    el,
    "hold the flag still while zooming (uncheck to zoom about the camera)",
    true,
    draw,
  );

  /** The last pointer position, in drawing space. Null until the reader moves over the canvas. */
  let pointer: { x: number; y: number } | null = null;

  const onMove = (event: PointerEvent) => {
    const box = canvas.getBoundingClientRect();
    // The canvas is scaled by CSS on a narrow screen, so displayed pixels are not drawing pixels.
    const drawn = parseFloat(canvas.style.width) || box.width;
    const k = box.width > 0 ? drawn / box.width : 1;
    pointer = {
      x: (event.clientX - box.left) * k,
      y: (event.clientY - box.top) * k,
    };
    draw();
  };
  const onLeave = () => {
    pointer = null;
    draw();
  };
  canvas.addEventListener("pointermove", onMove);
  canvas.addEventListener("pointerleave", onLeave);

  function params() {
    return {
      cameraX: cameraX(),
      cameraY: cameraY(),
      zoom: zoom(),
      anchorZoom: anchorZoom(),
    };
  }

  function polyline(points: Array<{ x: number; y: number }>, colour: string) {
    ctx.save();
    ctx.strokeStyle = colour;
    ctx.lineWidth = 2;
    ctx.beginPath();
    ctx.moveTo(points[0].x, points[0].y);
    for (const p of points.slice(1)) ctx.lineTo(p.x, p.y);
    ctx.stroke();
    ctx.restore();
  }

  function draw() {
    clear();
    const p = params();

    // Each layer is the same world drawn with a camera that has moved less. Far layers lag.
    LAYERS.forEach((layer, index) => {
      const cam = layerCamera(p, layer.factor);
      const at = (q: { x: number; y: number }) =>
        worldToScreen(cam, width, height, q);
      polyline(ridge(layer.factor, 1.6 + index * 0.5).map(at), layer.colour);
    });

    // The ground layer's own grid, so panning and zooming have something to be measured against.
    const ground = layerCamera(p, 1);
    for (let x = -24; x <= 24; x += 2) {
      const a = worldToScreen(ground, width, height, { x, y: -8 });
      const b = worldToScreen(ground, width, height, { x, y: 8 });
      line(ctx, a, b, GRID, { width: 1 });
    }
    for (let y = -8; y <= 8; y += 2) {
      const a = worldToScreen(ground, width, height, { x: -24, y });
      const b = worldToScreen(ground, width, height, { x: 24, y });
      line(ctx, a, b, GRID, { width: 1 });
    }

    // The flag: the landmark the anchor question is about.
    const flag = flagOnScreen(p, width, height);
    const flagBase = worldToScreen(ground, width, height, { x: FLAG.x, y: 0 });
    line(ctx, flagBase, flag, FLAG_COLOUR, { width: 2 });
    fillDot(ctx, flag.x, flag.y, 5, FLAG_COLOUR);
    label(ctx, "flag", flag.x + 8, flag.y - 4, FLAG_COLOUR);

    // The centre of the screen, which is where the camera is looking.
    line(
      ctx,
      { x: width / 2 - 7, y: height / 2 },
      { x: width / 2 + 7, y: height / 2 },
      TEXT,
      { width: 1 },
    );
    line(
      ctx,
      { x: width / 2, y: height / 2 - 7 },
      { x: width / 2, y: height / 2 + 7 },
      TEXT,
      { width: 1 },
    );

    // What the pointer is pointing at, converted back through the whole chain.
    const under = pointer ? worldUnderPixel(p, width, height, pointer) : null;
    if (pointer && under) {
      fillDot(ctx, pointer.x, pointer.y, 4, CURSOR);
      label(
        ctx,
        `world (${under.x.toFixed(2)}, ${under.y.toFixed(2)})`,
        pointer.x + 9,
        pointer.y - 6,
        CURSOR,
      );
    }

    const seen = visible(p, width, height);
    show(
      `camera at (${p.cameraX.toFixed(1)}, ${p.cameraY.toFixed(1)}) at ${p.zoom.toFixed(0)} px per unit \u00B7 ` +
        (seen
          ? `showing ${(seen.max.x - seen.min.x).toFixed(1)} by ${(seen.max.y - seen.min.y).toFixed(1)} units`
          : "nothing visible") +
        ` \u00B7 flag at pixel (${flag.x.toFixed(0)}, ${flag.y.toFixed(0)})`,
    );
    note(
      p.anchorZoom
        ? "sweep the zoom: the flag keeps its pixel, because the camera is moved to hold it there"
        : "sweep the zoom now: the flag slides away, because this zooms about the camera's own centre",
    );
  }

  draw();

  return () => {
    canvas.removeEventListener("pointermove", onMove);
    canvas.removeEventListener("pointerleave", onLeave);
  };
};

export default mount;

Pan with the first two sliders. The three ridges move at different rates because each is drawn with the same camera scaled down — more on that below. Move your pointer over the canvas and the world coordinate under it is read back through the whole chain.

Every click, tap and hover needs the other direction, and there is nothing new to derive:

pworld=(viewport⋅V)−1ppixelp_{\text{world}} = \left(\text{viewport} \cdot V\right)^{-1} p_{\text{pixel}}
A pixel to a world point and back, and the two ways to change zoom
The code src/lib/gamedev/demos/2d/screentoworld.ts
/** A pixel converted to a world point and back, and the two ways of changing zoom compared. */
import {
  camera,
  screenToWorld,
  unitsPerPixel,
  visibleWorld,
  worldToScreen,
  zoomAbout,
} from "../../../gamedev2d/camera2d.ts";
import type { Demo } from "../runner.ts";

const WIDTH = 620;
const HEIGHT = 330;
const at = (p: { x: number; y: number } | null) =>
  p === null ? "null" : `(${p.x.toFixed(2)}, ${p.y.toFixed(2)})`;

/** A camera looking at (8, 3) at 30 pixels per world unit. */
const CAM = camera({ x: 8, y: 3 }, 30);
/** The landmark the zoom is meant to hold still, several units off to the right. */
const FLAG = { x: 14, y: 5 };

const demo: Demo = (log) => {
  log(
    "camera at (8, 3), 30 px per unit, on a 620 by 330 canvas",
    `one pixel covers ${unitsPerPixel(CAM).toFixed(4)} world units`,
    "the number that decides whether art looks sharp",
  );
  log(
    "the middle pixel, converted to world",
    at(screenToWorld(CAM, WIDTH, HEIGHT, { x: WIDTH / 2, y: HEIGHT / 2 })),
    "the centre of the screen is exactly where the camera is",
  );
  log(
    "the top-left pixel, converted to world",
    at(screenToWorld(CAM, WIDTH, HEIGHT, { x: 0, y: 0 })),
    "left of the camera and above it, because the canvas counts y downward",
  );
  log(
    "world (14, 5) to a pixel and straight back",
    `${at(worldToScreen(CAM, WIDTH, HEIGHT, FLAG))} then ${at(
      screenToWorld(
        CAM,
        WIDTH,
        HEIGHT,
        worldToScreen(CAM, WIDTH, HEIGHT, FLAG),
      ),
    )}`,
    "the round trip every click depends on",
  );

  const seen = visibleWorld(CAM, WIDTH, HEIGHT)!;
  log(
    "how much world is on screen",
    `${(seen.max.x - seen.min.x).toFixed(2)} by ${(seen.max.y - seen.min.y).toFixed(2)} units`,
    "the canvas divided by the zoom, and the honest way to describe a zoom level",
  );

  // The two ways to change zoom, judged by what happens to the flag's pixel.
  const before = worldToScreen(CAM, WIDTH, HEIGHT, FLAG);
  const naive = { ...CAM, zoom: 60 };
  const anchored = zoomAbout(CAM, FLAG, 60);
  log(
    "zoom to 60 by assigning the number, and see where the flag went",
    `${at(before)} became ${at(worldToScreen(naive, WIDTH, HEIGHT, FLAG))}`,
    "it slid right off the canvas, because assigning the zoom zooms about the camera's centre",
  );
  log(
    "zoom to 60 about the flag instead",
    `${at(before)} became ${at(worldToScreen(anchored, WIDTH, HEIGHT, FLAG))}`,
    `unmoved, because the camera walked to ${at(anchored.position)} to keep it there`,
  );
};

export default demo;
camera at (8, 3), 30 px per unit, on a 620 by 330 canvas → one pixel covers 0.0333 world units // the number that decides whether art looks sharp
the middle pixel, converted to world → (8.00, 3.00) // the centre of the screen is exactly where the camera is
the top-left pixel, converted to world → (-2.33, 8.50) // left of the camera and above it, because the canvas counts y downward
world (14, 5) to a pixel and straight back → (490.00, 105.00) then (14.00, 5.00) // the round trip every click depends on
how much world is on screen → 20.67 by 11.00 units // the canvas divided by the zoom, and the honest way to describe a zoom level
zoom to 60 by assigning the number, and see where the flag went → (490.00, 105.00) became (670.00, 45.00) // it slid right off the canvas, because assigning the zoom zooms about the camera's centre
zoom to 60 about the flag instead → (490.00, 105.00) became (490.00, 105.00) // unmoved, because the camera walked to (11.00, 4.00) to keep it there

With the camera at (8,3)(8, 3) at 3030 pixels per unit, the middle pixel is exactly (8,3)(8, 3) — the camera’s own position always lands at the centre of the screen, which is the first thing to check when a camera misbehaves. The top-left pixel is (−2.33,8.50)(-2.33, 8.50): left of the camera and above it, because the canvas counts yy downward and the view matrix undid that.

Zoom About A Point, Or It Slides Out From Under You

Section titled “Zoom About A Point, Or It Slides Out From Under You”

Here is the bug everyone has felt. Change the zoom and leave the position alone, and you have zoomed about the camera’s centre — so whatever you were reaching toward slides away as you zoom in.

The fix is one line. Hold a world anchor’s view coordinates fixed while the zoom changes:

position′=a−zoomzoom′ (a−position)\text{position}' = a - \frac{\text{zoom}}{\text{zoom}'}\,(a - \text{position})

The camera has to walk toward the anchor. With the camera at (8,3)(8, 3), a flag at (14,5)(14, 5), and the zoom going from 3030 to 6060:

ApproachThe flag’s pixelThe camera ends at
assign the zoom(490,105)→(670,45)(490, 105) \rightarrow (670, 45)(8,3)(8, 3), unmoved
zoom about the flag(490,105)→(490,105)(490, 105) \rightarrow (490, 105)(11,4)(11, 4)

The first row slid the flag off a 620-pixel canvas entirely. The second did not move it by a hundredth of a pixel. And the camera’s new position, (11,4)(11, 4), is exactly halfway to the flag — because the zoom doubled, so the distance had to halve.

Try it in the scene above: uncheck the box and sweep the zoom. The rotation drops out of that formula, which is why it does not appear, and the build sweeps three anchors against twelve zoom pairs asserting the anchor holds its pixel — plus that the naive version genuinely moves it, so the assertion cannot pass by accident.

Distant things appear to move less. That is the entire effect, and it needs no new maths: draw the far layer with the same camera whose position has been scaled down.

FactorReads asBehaviour
11the actionmoves with the camera exactly
0.60.6middle treeslags a little
0.250.25far hillsbarely shifts
00the skynever moves at all

Only the translation is scaled. Touching the zoom too would make distant layers change size as you pan, which is not what distance looks like — so the check asserts zoom and rotation come through untouched, and that a factor of 0.250.25 shifts exactly a quarter as far as a factor of 11.

source A camera, both conversions, an anchored zoom, and parallax src/lib/gamedev2d/camera2d.ts 204 lines
/**
 * A camera, which is not a thing that looks at the world so much as a transform applied to it.
 *
 * The idea that makes cameras simple: a camera is **an object with a transform, used backwards**. Put
 * it at $(10, 4)$ and the world has to move $(-10, -4)$ to bring that spot to the middle of the screen.
 * Zoom in twice and the world scales by two, not by a half. So the matrix you actually draw with is the
 * **inverse** of the camera's own placement, which is why Section 3.2's inverse had to come first.
 *
 * Then the viewport: view coordinates have their origin in the middle of the screen with Y up, and a
 * canvas wants its origin top-left with Y down. That is Section 1.1's single conversion, and it lives
 * in exactly one function here.
 */
import {
  apply,
  compose,
  inverse,
  multiply,
  rotation,
  scaling,
  translation,
  type Mat3,
} from "./matrix2d.ts";
import type { Point } from "./vectors2d.ts";

/**
 * Where the camera is, how far in it is zoomed, and which way up it is.
 *
 * **`zoom` is pixels per world unit.** At 30, a one-unit tile is 30 pixels across and a 620 pixel
 * canvas shows about 20 units of world. Doubling it to 60 draws everything twice as big and shows half
 * as much, which is what a reader means by zooming in.
 *
 * Stating it in pixels per unit rather than as an abstract multiplier is worth the extra word: it is
 * the number that decides whether art looks sharp, it is what `unitsPerPixel` inverts, and it means
 * "what zoom should this be" has an answer you can work out from your tile size instead of guess.
 */
export type Camera = {
  position: Point;
  zoom: number;
  /** Radians. Counter-clockwise in world terms, as everywhere else in this module. */
  rotation: number;
};

export function camera(
  position: Point = { x: 0, y: 0 },
  zoom = 30,
  rotationRadians = 0,
): Camera {
  return { position, zoom, rotation: rotationRadians };
}

/** Zoom must stay positive: zero would flatten the world to a point and negative would mirror it. */
export function withZoom(cam: Camera, zoom: number, min = 1e-3): Camera {
  return { ...cam, zoom: Math.max(min, zoom) };
}

/**
 * The camera treated as an ordinary object in the world, which is the thing being inverted.
 *
 * Note the scale is $1/\text{zoom}$. A camera zoomed in twice is a small window on the world, so as an
 * object it is half the size - and inverting that is what makes the world twice as big.
 */
export function cameraPlacement(cam: Camera): Mat3 {
  return compose(
    translation(cam.position.x, cam.position.y),
    rotation(cam.rotation),
    scaling(1 / cam.zoom, 1 / cam.zoom),
  );
}

/**
 * World to view: the view matrix. Every term is the opposite of the camera's own.
 *
 * $$V = S(\text{zoom}) \; R(-\theta) \; T(-\text{position})$$
 *
 * Read right to left, as Section 3.1 requires: move the world so the camera sits at the origin,
 * un-rotate it, then scale it up by the zoom. Written out this way it needs no matrix inversion at
 * runtime, and the build checks it really is `inverse(cameraPlacement)` rather than merely resembling
 * it - three sign errors would each still produce a believable picture.
 */
export function viewMatrix(cam: Camera): Mat3 {
  return compose(
    scaling(cam.zoom, cam.zoom),
    rotation(-cam.rotation),
    translation(-cam.position.x, -cam.position.y),
  );
}

/**
 * View to pixels: put the origin in the middle of the canvas and flip Y.
 *
 * The whole of Section 1.1 in one matrix, and the only place in this file that knows a screen exists.
 * Keeping it separate is what lets the same camera drive a phone and a monitor.
 */
export function viewportMatrix(pixelWidth: number, pixelHeight: number): Mat3 {
  return multiply(translation(pixelWidth / 2, pixelHeight / 2), scaling(1, -1));
}

/** The full chain, world all the way to pixels. One matrix, built once per frame. */
export function worldToScreenMatrix(
  cam: Camera,
  pixelWidth: number,
  pixelHeight: number,
): Mat3 {
  return multiply(viewportMatrix(pixelWidth, pixelHeight), viewMatrix(cam));
}

export function worldToScreen(
  cam: Camera,
  pixelWidth: number,
  pixelHeight: number,
  p: Point,
): Point {
  return apply(worldToScreenMatrix(cam, pixelWidth, pixelHeight), p);
}

/**
 * Pixels back to world. What every click, tap and hover needs.
 *
 * Just the inverse of the chain above, which is the point: there is nothing to derive separately and
 * nothing to keep in step. Deriving a second formula by hand is how the two drift apart, and a
 * screen-to-world that is slightly wrong still returns a plausible position - so this is the function
 * whose round trip the build sweeps over the whole canvas.
 */
export function screenToWorld(
  cam: Camera,
  pixelWidth: number,
  pixelHeight: number,
  pixel: Point,
): Point | null {
  const back = inverse(worldToScreenMatrix(cam, pixelWidth, pixelHeight));
  return back === null ? null : apply(back, pixel);
}

/**
 * Change the zoom while holding one world point still on screen.
 *
 * $$\text{position}' = a - \frac{\text{zoom}}{\text{zoom}'}\,(a - \text{position})$$
 *
 * The default, and wrong, way to zoom is to change the number and leave the position alone. That zooms
 * about the **camera's centre**, so whatever you were looking at slides away as you zoom toward it -
 * which every reader has felt in a map that zooms out from under their finger.
 *
 * Solving for it is short. The anchor's view coordinates are $\text{zoom} \cdot R^{-1}(a - \text{pos})$,
 * so holding them fixed while zoom changes forces the displacement to shrink by exactly the ratio of
 * the two zooms. The rotation drops out, which is why it does not appear above.
 */
export function zoomAbout(cam: Camera, anchor: Point, newZoom: number): Camera {
  const next = Math.max(1e-3, newZoom);
  const k = cam.zoom / next;
  return {
    ...cam,
    zoom: next,
    position: {
      x: anchor.x - k * (anchor.x - cam.position.x),
      y: anchor.y - k * (anchor.y - cam.position.y),
    },
  };
}

/**
 * The camera a parallax layer should be drawn with: the same camera, moved less.
 *
 * A factor of 1 is the layer the action happens on. Below 1 the layer lags behind, which reads as
 * distance; a factor of 0 never moves at all, which is a sky. **Only the translation is scaled** -
 * touching the zoom as well would make distant layers change size as you pan, which is not what
 * distance looks like.
 */
export function parallax(cam: Camera, factor: number): Camera {
  return {
    ...cam,
    position: { x: cam.position.x * factor, y: cam.position.y * factor },
  };
}

/**
 * The axis-aligned world rectangle currently on screen, for culling and for knowing what to load.
 *
 * Taken as the bounding box of the four unprojected corners, so it stays correct under a rotated
 * camera - where the visible region is a tilted rectangle and its bounding box is deliberately larger.
 */
export function visibleWorld(
  cam: Camera,
  pixelWidth: number,
  pixelHeight: number,
): { min: Point; max: Point } | null {
  const corners = [
    { x: 0, y: 0 },
    { x: pixelWidth, y: 0 },
    { x: pixelWidth, y: pixelHeight },
    { x: 0, y: pixelHeight },
  ].map((c) => screenToWorld(cam, pixelWidth, pixelHeight, c));
  if (corners.some((c) => c === null)) return null;
  const xs = corners.map((c) => c!.x);
  const ys = corners.map((c) => c!.y);
  return {
    min: { x: Math.min(...xs), y: Math.min(...ys) },
    max: { x: Math.max(...xs), y: Math.max(...ys) },
  };
}

/** How many world units one pixel covers. The number that decides whether a sprite looks sharp. */
export function unitsPerPixel(cam: Camera): number {
  return 1 / cam.zoom;
}
  • Every click and tap in a game with a moving camera.
  • Pinch and scroll zoom, which is the anchored formula with the anchor at the gesture’s centre.
  • Culling, using the visible world rectangle to skip what is off screen.
  • Parallax backgrounds, and minimaps, which are a second camera on the same world.
  • Section 4.1, where the camera stops snapping to a target and starts following it — and needs frame-rate independence to do that without jitter.
  • The capstone’s camera, which follows a character without shaking.