Philosophy That Works

2026-09-30 · Jin Ha

A seat map for the web

Can every AI read a web page the same way, and know exactly where to click? We gave the web a seat map.

1Question

It started with one question: could any AI read a web page in the same way?

A person looks at a shopping site and taps the cart in the top right corner without thinking. For an AI, that is hard. An AI usually reads a web page in one of two ways: it looks at a screenshot, or it reads only the text. A screenshot is easy to misread. The text alone doesn't say where anything is on the screen.

But for an AI to click or type for you, it has to know exactly where the mouse and keyboard should go.

2Structure

So we drew a grid over the page, like a game board. Twelve columns across, and a number for every row down. Everything you can press gets a number tag, along with the row and column where it sits.

It works like a theater seat map. Say "row 3, seat 5" and anyone finds the same seat. Grid makes a seat map like this for each web page and hands it to the AI. It also tells the AI what is happening right now: this button is covered by a pop-up, this button can't be pressed yet.

We call this way of writing the seat map GRID/1.

3Experiment

We gave four AI models tasks on real web pages. Once, they got only a list of what was on the page, without positions. Once, they got the Grid seat map. Then we checked whether they found the right thing.

  • Ordinary pages: with the list alone, they got 87% right. With the seat map, 99.5%.
  • Trap pages, where the order of the text differs from where things appear on screen: with the list alone, they fell to 61%. With the seat map, all four models got 100%.

For example, we asked: "Put two red T-shirts, size M, in my cart." The seat map told the AI three things: a cookie notice was covering the buy button, so button 14, "OK", had to be pressed first; color and size had to be chosen; and "Add to cart" couldn't be pressed yet. Every model planned the same steps: 14 "OK" → red → size M → quantity 2 → Add to cart.

4Understanding

However smart a model is, if it misreads the page, it cannot understand it. But with a seat map, we raised its understanding.

Grid is AI's new pair of eyes. For now it reads; the hands that click come next. Grid runs on its own, so any AI can call it as a tool, and it is becoming the eyes of our own product, ieum.

The next questions are already here. How do we read words inside a picture? How do we read a page where video, pictures, and words are mixed together?