JavaEye

Why Your Mobile PageSpeed Insights Score Is Hard to Move — and How I Pushed Five Pages From 83 to 99–100

The main bottleneck on my static tool site was not my code. It was a single 177KB third-party analytics script, and moving its load out of the parse window took five pages from 83–96 to 99–100 on Google PageSpeed Insights (mobile), with LCP dropping from 3.5s to 1.5s. This post is the full case study: the measurement traps, the numbers, the code, and the second failure I did not see coming.

Quick numbers (TallyPunch, official PSI mobile, September 2026):

Measurement behavior at PSI changes over time; treat the caching quirks below as “verify against current docs,” not eternal law.

Set your baseline on official PageSpeed Insights, not local Lighthouse

If you have ever wondered why your mobile PageSpeed Insights score is so hard to move while local Lighthouse says everything is fine — the two tools are measuring different worlds. The same page, at the same moment:

MeasurementScores
Local Lighthouse (my network, mainland China)39 / 63 / 57
Official PageSpeed Insights (Google’s servers)96 (three runs, identical)

Local Lighthouse is the open-source audit engine running on my machine over my network. PageSpeed Insights runs Lighthouse on Google’s hardware over a simulated mobile network, blended with real-user data. On my restricted network, the third-party script’s download and execution timing lands randomly inside or outside the measurement window, and the local variance becomes too large to optimize against. Google’s official Core Web Vitals thresholds — LCP 2.5s, INP 200ms, CLS 0.1 — are the same numbers PSI scores against, so a baseline taken anywhere else is optimizing for the wrong target.

One more trap: PSI serves cached responses for identical URL + strategy requests. Three “consistent” reads may be one cached response. Break the cache before trusting the median, and compare against an untouched control page from the same session.

The failing resource is usually not your own code

My site’s entire bill, from the PSI third-parties-insight audit, looked like this:

ItemMy codeThird parties
Transfer bytesHTML + CSS + favicon ≈ 14KBgtag.js 177KB + CF beacon 10KB
Share of 194KB total~7%93%
Main-thread time~90msgtag 171ms (786ms blocking, local)
CLS0.000—

Fourteen kilobytes of my own code, zero forced reflows, zero layout shift. Nothing to fix there — which was almost disappointing, because I had a profiler ready and no villain of my own to catch.

The diagnosis lives in three PSI audits, not in the headline score: third-parties-insight shows who eats the main thread, total-byte-weight shows who eats bytes, and long-tasks shows the causal chain — in my case a long task starting at 2429ms against an LCP of 2551ms, the task sitting exactly on top of the paint moment. That is causation, not correlation.

This also falsified, for this site shape, the common claim that TTFB is the upstream cause of LCP. My server-response-time was 1ms. On a small static site, the bottleneck is bytes and main-thread time, not the server. The TTFB chain holds or breaks depending on what kind of site you run — quantify before optimizing.

Local A/B tests prove the mechanism, never the magnitude

A controlled localhost A/B — same build, one variable, interleaved runs — cleanly isolated the mechanism. gtag static tag versus deferred load: TBT 343ms versus 0ms. Mechanism confirmed.

But two structural blind spots make local numbers non-portable. First, the RTT blind spot: localhost has effectively zero round-trip time, while PSI’s simulation charges 150ms per request. Both local variants capped at 99 with unchanged LCP; the same page scored 83 on PSI, and the entire gap lived in the RTT chain. “No change locally” does not mean “no change on PSI.” Second, my restricted network inflated the third-party cost about fivefold — 786ms of blocking locally versus 171ms in PSI’s environment. Direction agreed, magnitude did not.

The rule I kept: local A/B answers “does the mechanism work,” and official PSI answers “how much does it buy.” Mixing the two wastes an afternoon, and I speak from experience.

Defer third-party scripts until first interaction — and pin the delay with a test

The fix keeps analytics semantics completely intact; it only moves the library’s injection later:

<script>
  window.dataLayer = window.dataLayer || [];
  function gtag(){dataLayer.push(arguments);}
  gtag('js', new Date());
  gtag('config', 'G-XXXXXXX');          // config stays inline — semantics unchanged
  (function () {
    var src = 'https://www.googletagmanager.com/gtag/js?id=G-XXXXXXX';
    var FALLBACK_MS = 10000;            // named constant, locked by a unit test
    var done = false;
    function load() { if (done) return; done = true;
      var s = document.createElement('script'); s.async = true; s.src = src;
      document.head.appendChild(s); }
    ['pointerdown','keydown','touchstart','scroll'].forEach(function (ev) {
      window.addEventListener(ev, load, { once: true, passive: true });   // real users: no data lost
    });
    window.addEventListener('load', function () {
      setTimeout(function () {                                            // passive visits: out of the window
        if (window.requestIdleCallback) requestIdleCallback(load, { timeout: 2000 });
        else load();
      }, FALLBACK_MS);
    }, { once: true });
  })();
</script>

Three load-bearing details. The dataLayer stub and config call stay inline in the document, so no tracking call site changes. Real users trigger the load on first interaction — click, key, touch, or scroll — so session data survives. The fallback fires 10 seconds after load for passive visitors, pushing the script out of the lab measurement window.

The honest cost: a visitor who reads passively and leaves within ten seconds is not counted. That trade was worth it for my content site, but it is a real trade, and I keep the books on it.

The second failure mode is subtler: setting FALLBACK_MS too low. A two-second delay drags the long task straight back into the TBT measurement window — the score collapses by points, and six months later nobody remembers why. My unit test asserts FALLBACK_MS >= 8000, bans any static <script async src=...> tag, and checks all four interaction triggers exist. The constant has a name because the number is a promise.

A site-wide fix patched into one page is debt, not optimization

Then I audited a second site — TypeOnSkin, an interactive tattoo-font editor — and the main cause was the opposite of everything above. Not third parties. My own first-screen CSS, linked externally, render-blocking: main.css wasted 151ms and fonts.css wasted 451ms, and the external stylesheet delayed the font requests until CSS finished parsing. The fix script that inlined critical CSS only served the homepage.

The homepage had already been cured of this exact disease one day earlier. The test that guards it even carried the diagnosis in its assertion message — “head must have no render-blocking external CSS (v1 externals caused mobile PSI 79)” — but that assertion only scanned index.html. The cause of death was written inside the guardrail, and the guardrail watched one page. Every new page type re-caught the disease, legally, with nobody remembering the original outbreak.

The general form, and the most valuable sentence in this post: a site-wide optimization implemented as a patch on one page is not an optimization — it is debt. Every new page type re-runs it, and the rerun never remembers why.

The repair had three parts, all required: inline critical CSS in the build script’s shell itself rather than as a post-build injection step; one shared inlineStyleTag() constructor used by both baking and post-processing; and a whole-site invariant, asserted across every HTML file, that the head contains zero rel="stylesheet" links. Structure beats score for verifying recovery — scores get noisy, structure does not. All 24 pages now ship zero render-blocking stylesheets and zero CSS requests. (The font preload I almost added would have been negative value: the LCP element is system-ui body text, so preloading fonts only competes for bandwidth. Check what your LCP element actually is before optimizing for it.)

How to improve Core Web Vitals: the ten steps, in order

This is the checklist my next site will run from day one — a Core Web Vitals case study compressed into mechanics:

  1. Take the baseline on official PSI. Break its cache first, run five or more passes, take the median, and judge against an untouched control page from the same session.
  2. Open third-parties-insight and total-byte-weight. Identify who owns the bytes and the main thread before touching anything.
  3. Compare long-tasks start times against LCP. A long task starting at the paint moment is the causal chain you are looking for.
  4. Audit your own first-screen CSS too. Look for external rel="stylesheet" in the head and wasted milliseconds in render-blocking-resources — the second most likely villain.
  5. Defer third-party scripts until first interaction, with an idle fallback well outside the measurement window.
  6. Lock the deferral in tests: no static script tags, a named FALLBACK_MS constant with a floor assertion, all interaction triggers present.
  7. Build site-wide invariants into the build itself, with one shared constructor and whole-site assertions — never a per-page patch.
  8. Identify the LCP element before any font or image strategy. Preloading assets the LCP element does not use is negative value.
  9. Re-measure on PSI after every change. Local numbers confirm mechanism only.
  10. Collect the side findings — my pass also surfaced 16 color-contrast failures and security-header suggestions. The audit you ran anyway owes you the small print.

The results, before and after (official PSI, mobile)

PageScoreLCPTBT
Home96 → 1002.55s → 1.52s100ms → 0ms
time-card83 → 993.52s → 1.52s0 → 0ms
biweekly83 → 993.53s → 1.53s0 → 0ms
double-time90 → 993.49s → 1.55s45ms → 0ms
decimal90 → 1003.51s → 1.53s60ms → 0ms

The site is TallyPunch, a free no-login time-card calculator suite for US small businesses. The 93%-third-party pattern should generalize to any small static site running analytics; the render-blocking relapse warns anyone maintaining more than one page type.


This article was created with the help of AI. AI was not used to write the content; it assisted only with translation and grammar checks.

This article was created with the help of AI