The Eye Doesn't Read: Hierarchy and Gestalt
In the first seconds nobody reads your screen — they scan it for something that looks like the answer. Design decides what that something is.
You open a page to find out what the thing costs. Your eyes jump: big picture, some heading, a block of grey text ignored entirely, a box on the right, back to the middle, down, there it is. Four seconds. You read perhaps eleven words.
Now recall how that page was made. People argued about the wording of a paragraph you never looked at, and nobody decided which single element you would see first. That's backwards, and this article is about putting it the right way round.
Three seconds of reconnaissance
The evidence on this is old and boring, which is the best kind. Studies of real browsing sessions — Weinreich and colleagues in 2008, and Jakob Nielsen's analyses before them — put the share of words actually read on a typical page at something like a fifth to a quarter. Not a quarter of the sentences: a quarter of the words, scattered.
Two more findings from eye-tracking are worth carrying:
The F-pattern. On text-heavy pages with weak structure, gaze traces a rough F: across the top, a shorter sweep lower down, then a slide along the left edge. It isn't a law of nature — it's what people fall back on when nothing on the page tells them where to look. Strong hierarchy overrides it. Its real message is: the first two words of every line and every heading do most of the work.
Banner blindness. Benway and Lane described it in 1998: people don't just ignore ads, they become unable to see anything that resembles one. Bright, boxed, off to the right, slightly decorative — the eye filters it before consciousness arrives. Plenty of important product messages have been made invisible by being made to look important in the wrong way.
Underneath both sits a behaviour Steve Krug borrowed from Herbert Simon: satisficing. People don't survey the options and choose the best. They take the first thing that looks like it might work, because backing out is cheap and reading everything is not.
They don't pick the best option. They pick the first plausible one.
Which means being available beats being complete. The right answer in paragraph three loses to a decent answer in the heading.
What the eye sees before you do
Some visual differences are processed in parallel, before attention arrives — under about 200 milliseconds, and at essentially the same speed whether there are ten objects or a hundred. Anne Treisman's work on feature integration mapped them. The practical list is short: colour, size, orientation, motion, enclosure, position.
Anything on that list makes an element findable instantly. Anything not on it — an icon shape, a word, a subtly different font — requires a serial hunt, item by item.
There's a catch, and it's the one people miss. The effect works only when the difference is singular. One red item among grey ones is found instantly. Ten red items among twelve grey ones are just a mess with no signal at all. Hedwig von Restorff demonstrated the memory version of this in 1933: the item that differs from its neighbours is the one recalled — but only while it is the one that differs.
So emphasis is a budget, not a property. Every element you make louder devalues every other loud element on the screen. A designer with five priorities has none.
Gestalt: how the eye decides what belongs together
Before a person understands anything on your screen, they've already grouped it — automatically, in milliseconds, according to principles the Gestalt psychologists described in the 1920s. You aren't choosing whether grouping happens. You're choosing whether it groups the things you meant.
Proximity. Objects close together read as one group. This is the strongest of the lot and beats almost anything else — including colour and similarity. It's also free: space costs nothing to add.
Similarity. Things that share shape, colour or size read as belonging to the same class. This is what makes "all buttons look like buttons" a functional rule and not a stylistic one.
Common region. Anything inside a shared boundary — a card, a panel, a tinted background — reads as one group, even overriding proximity. That's why a card is the workhorse of modern interfaces.
Continuity. The eye follows lines and alignments and expects them to continue. Aligned elements read as a sequence; a misaligned one reads as unrelated, or as a mistake.
Closure. People complete incomplete shapes. A cut-off row of items at the edge of the screen tells the eye "there is more here, scroll" — the most useful accidental affordance in interface design.
Common fate. Things that move together belong together. Animation isn't decoration; it's a grouping statement.
The classic bug produced by ignoring all of this: a form where each label sits directly above its field but with generous space, and the fields are tight together. Proximity groups every label with the field below it. People fill in the wrong boxes and nobody can say why. The fix is not a better font — it's four pixels moved.
Hierarchy is an answer, not a style
Visual hierarchy sounds like an aesthetic concept. It isn't. It's the answer to one question: in what order will this be seen? You either answer it or the layout answers it for you, arbitrarily.
Five tools do the work, roughly in order of strength: size, weight, contrast, space, position. Colour appears in that list only as contrast — in a monochrome interface, hierarchy loses nothing.
Two working rules:
Three levels per screen, maximum. Primary — the one thing this screen is for. Secondary — supporting information, findable when sought. Tertiary — everything else, deliberately quiet. Four levels are indistinguishable in practice, and everyone tries anyway.
Every screen has exactly one primary. If two things are fighting to be first, the decision hasn't been made yet — it's been passed on to the user, who will resolve it by leaving.
The cheapest check ever invented: squint at your screen. Blur your vision until only shapes remain. Whatever is still visible is your real hierarchy. If that isn't what you'd want a stranger to see first, nothing else on the screen matters yet.
Empty space is doing work
There is constant pressure to fill space — from stakeholders who see emptiness as waste, and from a sincere belief that more visible content means more use.
But space is not absence; it's the primary grouping tool, and grouping is most of what a person is doing in those first three seconds. Removing space doesn't add information — it destroys the structure that made the information findable.
The "above the fold" argument is the usual weapon here, and it's largely a myth carried over from newspapers. People scroll, and have done for two decades. What actually kills scrolling is a screen whose bottom edge looks like an ending: a full-width band, a neat closing line, nothing cut off. The Gestalt fix is closure — let something be visibly interrupted by the bottom of the screen, and the eye knows to continue.
In practice
The squint test. After every screen you make, blur your eyes and name what you see first, second, third. Do this before showing anyone anything.
The greyscale test. Turn the design black and white. Everything built on colour alone collapses, and you find out what your hierarchy actually rests on. This is also the fastest accessibility check you can run in five seconds.
Measure your gaps, don't eyeball them. For any group of elements, the space inside the group must be visibly smaller than the space around it. Not slightly — visibly. Most "cluttered" interfaces are one pass of this away from being fine.
Count your emphases. Go over a screen and count everything shouting: bold text, coloured elements, boxes, big type. More than three and you've spent the budget. Demote, don't add.
Cut something off at the fold. Make sure the bottom of the first screen visibly continues. A half-visible card is worth more than any "scroll down" arrow.
Five-second test. Show your screen to someone for five seconds, take it away, and ask what it's for and what they'd do next. It's the cheapest real test in this sphere, and you can run it on a colleague in a corridor.
Check yourself
Close the article and answer in your own words:
- Roughly what share of words on a page gets read, and what does that imply about where you put the answer?
- What does the F-pattern actually tell you — a rule about eyes, or a symptom of the page?
- Why does making five things stand out mean nothing stands out?
- Which Gestalt principle is strongest, and which one can override it?
- How does bad spacing make people fill in the wrong form field?
- What is the "one primary per screen" rule really about?
- What makes a page look unscrollable, and which principle fixes it?
In short
- People read a fraction of the words and satisfice: they choose the first plausible option, not the best one.
- The eye finds colour, size, orientation, motion and enclosure pre-attentively — but only when the difference is singular. Emphasis is a budget.
- Gestalt principles group elements automatically: proximity, similarity, common region, continuity, closure, common fate. You don't choose whether — only whether it groups what you meant.
- Proximity is the strongest and the cheapest tool you have. Most clutter is a spacing problem, not a content problem.
- Hierarchy answers "what is seen first." Three levels maximum, one primary per screen, built from size, weight, contrast, space and position.
- Space is a tool, not waste. The fold is a myth; a screen that looks finished at the bottom is what actually stops scrolling.
- Squint test, greyscale test, five-second test — three checks that cost nothing and catch most of it.