Skip to main content
    ← lab / researchR-01 · PRE-REGISTERED PROTOCOL + DESK STUDY

    The five-second gap

    Can an AI stand in for a human first impression? This autumn, Blu — our resident AI — takes the classic five-second test against real people, on the same 40 landing pages. Before we collect a single data point, here is everything we've locked in: why five seconds, what the science already says, and exactly how we'll score the gap.

    blurple lab researchprotocol v1.0 — registered 28 Jul 20266 min readFIELDWORK: AUTUMN 2026
    fig. 1 — a five-second test, sliced into 50-millisecond ticks. By the first teal tick, a visitor's visual-appeal judgment has already formed (Lindgaard et al., 2006). The other 99 ticks decide what they think you do.

    Every founder believes their homepage is clear. The five-second test — flash the page, take it away, ask what it said — has been puncturing that belief for two decades, because it measures the only audience that matters: someone who owes you nothing. Panels are slow and cost money. A vision model is instant and free. Whether it can honestly do the same job is an open question, and we'd rather measure it than market it.

    50 ms
    is all it takes to form a visual-appeal judgment of a web page — and it correlates strongly with judgments made at longer exposures. (Lindgaard et al., 2006)
    46.1%
    of users assess a site's credibility primarily by its visual design — not its content. (Stanford Web Credibility Project, 2002)
    200
    human five-second sessions in our fieldwork: 40 landing pages × 5 participants each, against Blu's verdicts on the identical set.
    01 — WHY FIVE SECONDS WORKS

    The first impression is a verdict, not a glance.

    Lindgaard's landmark 2006 studies flashed homepages at people for half a second, then a twentieth of a second — and got the same appeal ratings both times. The judgment is visceral, it forms before reading begins, and it doesn't stay in its lane: through the halo effect, that instant aesthetic verdict bleeds into how usable and how credible people believe the site is. Stanford's web-credibility research found nearly half of users lean on visual design when deciding whether to trust a site at all.

    The five-second test, popularized by usability researchers in the mid-2000s, is the practical wrapper around this science: five seconds is long enough to read a headline and notice one or two elements, short enough that nothing can be studied. What survives those five seconds is your actual message — everything else is what you hoped your message was.

    02 — THE OPEN QUESTION

    Where we expect Blu to agree with humans — and where not.

    Our tool blu first-glance already runs AI five-second tests. Honesty requires the follow-up: has anyone checked the AI against the humans it imitates? We registered three hypotheses before touching data:

    H1High agreement on attention — what gets noticed first. Saliency is the thing vision models are literally trained on.
    H2Lower agreement on comprehension — "what does this company do?" Humans misread in human ways; models misread in confident ways.
    H3The gap widens off the beaten path — non-English pages and unconventional layouts. If true, this defines exactly where AI testing needs a human in the loop.

    If Blu fails, we publish the failure and blu first-glance gets a printed warning label. If Blu holds up, every founder gets a five-second panel for free. Either result is worth having — which is the test of a question worth asking.

    03 — THE PROTOCOL, LOCKED

    Registered before the first data point.

    This page is the pre-registration. The design below was frozen on 28 July 2026; any deviation will be listed in the published results.

    SAMPLE40 SaaS landing pages: 20 English, 10 Turkish, 10 Spanish — stratified by traffic rank, screenshotted on the same day at two viewports.
    HUMAN ARM5 participants per page (200 sessions) via a moderated five-second protocol. Three questions, verbatim: What does this company do? What did you notice first? Where would you click next?
    AI ARMBlu first-glance on the identical screenshots, same three questions, fixed prompt and temperature, 3 runs per page to measure the model's own consistency.
    SCORINGTwo coders, blind to source, match human and AI answers per question; agreement reported as Cohen's κ per question class, per language.
    OUTPUTAgreement matrix, full transcripts, prompts and raw CSV — published open, here. Argue with us: hello@blurple.digital
    An honest note on what this can't say.

    A five-second test measures first impressions — not brand memory, market context, or whether anyone converts on day thirty. And 200 sessions is a field study, not a census: enough to size the gap, not to end the argument. That's why the raw data ships with the result.

    sources
    [1] Lindgaard, Fernandes, Dudek & Brown (2006) — "Attention web designers: You have 50 milliseconds to make a good first impression!" Behaviour & Information Technology, 25(2), 115–126
    [2] Fogg et al., Stanford Web Credibility Project (2002) — how users assess website credibility
    [3] UIE — the five-second test method
    [4] Nielsen Norman Group — the aesthetic-usability effect
    blurple lab · research
    all studies →
    © 2026 blurple studio · Certified B Corporationmethod, prompts & raw data ship with every study