Log inPreview your site
Sampled: ChatGPT · Gemini · monthlyReceipts kept · unsampled ≠ absentYou approve every change — verified live

measurement · field guide

How to read your free Foliora site snapshot

What a bounded public scan can establish deterministically, what remains a sample of the strategy shape, and what connected Foliora research adds next.

The snapshot is intentionally bounded

The planned public preview reads roughly 10–20 canonical public pages. It does not sign in, bypass controls, inspect private data, or claim to represent every URL on the site.

Canonical public pages means the pages a visitor — or a crawler — would actually land on without special access: the homepage, the primary navigation destinations, and the pages internal links treat as authoritative, not every parameter variant, paginated listing, or duplicate URL a larger site accumulates. The bound exists because a public preview has to stay fast and non-invasive; a full crawl of every URL on a site the size most prospects run would take considerably longer and would need permissions a public tool shouldn't be asking for before someone has decided to sign up.

That bound has a direct consequence for how to read the result. A defect found on several of the sampled pages — a missing canonical, a templated title repeated with no real variation — is strong evidence of a sitewide pattern, because the same template usually produces the same defect everywhere it's used. A defect not found in the sample isn't the same kind of evidence in reverse: it means those specific pages looked fine, not that every URL on the site does. The snapshot is a representative look at a site's technical foundation, not an audit of every page on it.

From a bounded scan to connected research

Each stage answers a different kind of question, with a different kind of evidence behind it.

1

Bounded scan

Reads roughly 10–20 canonical public pages — no login, no bypassing access controls, no private data.

2

Deterministic findings

Titles, headings, canonicals, response behavior, links, structured relationships, and source signals — observable, checkable facts.

3

Sampled opportunities

Buyer questions and strategy shape drawn from what's visible — a preview of the shape, not a full market or competitor study.

4

Connected deep research

Full demand, competitor, source, crawl, and citation research after connecting an account and a site.

Why these particular checks, and not others

Titles, canonicals, robots rules, sitemaps, and structured data aren't an arbitrary shortlist — they're close to the minimal set of signals that are both fully within a site owner's control and directly documented by the search engines that read them. Google is explicit that a robots.txt file's job is managing what a crawler is allowed to request, not what ends up indexed, and that a page blocked there can still surface in results without a description if it's linked to from elsewhere — which is exactly why robots.txt is worth checking directly rather than assuming a disallow rule does more than it actually does.

Canonicals matter for the same reason a scan checks them among the first on-page signals: Google's own canonicalization documentation describes a declared canonical as a hint it weighs against redirects, sitemap inclusion, and its own signals like HTTPS preference — not a rule it's obligated to follow. A site that gets that hint internally inconsistent (one canonical in the markup, a different URL in the sitemap) gives Google more reason to override the stated preference than a site that's internally consistent, even before ranking enters the conversation at all.

Titles are checked because Google's documentation names the exact patterns that make it discard a site's own title tag in favor of something it generates instead — a title that doesn't match the page's main visible heading, a boilerplate pattern repeated across a page series, or no single element that clearly reads as the page's main title. And structured data is checked against two separate bars Google documents: a technical bar (valid format, required properties present) that a validator catches, and a quality bar (the markup actually describes what's visible on the page) that no validator catches and that determines whether the markup helps or creates a policy problem instead.

Foundational findings are deterministic

The preview can inspect visible titles, headings, canonicals, response behavior, links, structured relationships, and source signals. These checks establish observable conditions before any model summarizes them.

Each of those checks exists because it maps to something search engines and assistants actually use to decide whether a page can be found, trusted as the authoritative version, and represented accurately. A title and heading that disagree in substance is one of the documented reasons Google rewrites what it shows in a search result instead of using the page's own title tag. A missing or duplicated canonical leaves a search engine to guess which of several similar URLs is the one worth showing, and its guess is a hint-following process, not a guarantee it lands on the page a site owner intended. Response behavior — status codes, redirect chains, whether a URL that should be gone actually returns a 404 instead of quietly redirecting to the homepage — determines whether a crawler's limited attention gets spent efficiently or wasted retracing paths that don't lead anywhere new.

Links and structured relationships are checked because discovery depends on them independent of anything a sitemap claims: a page with no real anchor-tag link pointing to it from an already-crawled page is functionally orphaned, no matter how prominent it looks in a site's navigation design. Source signals — what a site's robots.txt and sitemap actually say, as opposed to what a site owner assumes they say — round out the deterministic set, because both files are simple enough to be either exactly right or subtly wrong in a way that's easy to miss without checking the live file directly rather than the intent behind it.

None of these checks require interpretation to be useful. A canonical tag either resolves to a 200 on the intended URL or it doesn't; a title element either matches the page's main heading in substance or it doesn't. That's the sense in which these findings are deterministic rather than sampled — they're facts about what a site is currently serving to any crawler, verifiable independently by rerunning the same check, not an estimate of likely performance.

Reading a finding: what it means, and what it doesn't

A deterministic finding — a missing canonical, a duplicated title, a robots rule that blocks more than it should — describes a condition on the live site right now. It doesn't, on its own, describe how Google has already responded to that condition on a given URL, because that requires access to something a public scan structurally doesn't have: the actual Search Console indexing status for that property.

That's a meaningful limit worth stating plainly. Google draws a hard line between "Discovered — currently not indexed" (found but not yet crawled) and "Crawled — currently not indexed" (crawled, then deliberately not added to the index) — and telling those two apart for a specific site requires that site's own Search Console data, not an outside scan. A public snapshot can identify the kind of technical condition that tends to produce either outcome — a thin or duplicated page, a canonicalization conflict, an orphaned URL — but it can't tell a site owner which URLs are actually sitting in which status without a connected Search Console property to check against.

The practical takeaway: treat a deterministic finding as a defect worth fixing on its own terms — a missing canonical is wrong regardless of what it's currently doing to any specific ranking — rather than as a claim about current search performance. Confirming the actual before-and-after effect on indexing or visibility is exactly the kind of question a connected account and its own Search Console data are positioned to answer, and the free snapshot doesn't claim to answer it in their place.

Buyer questions and opportunities are samples

The free result shows enough of the strategy shape to be useful, not a comprehensive market study. Full demand, competitor, source, crawl, and citation research requires an account and connected site.

The distinction is about what kind of evidence backs each part of the result. The deterministic findings above are checkable directly against the live site — anyone could rerun them and get the same answer. Buyer questions and opportunity framing are drawn from what's visible in the sampled pages' content and structure, which is a reasonable starting signal for the shape of a strategy but isn't the same as measured demand, tracked competitor behavior, or observed citation results. Those require data a bounded public scan structurally can't have: actual search query volume and click behavior, which lives in Search Console tied to a specific verified property; a competitor's actual content and technical setup beyond what's publicly visible; and how AI assistants actually respond to real prompts over time, rather than a single bounded probe.

This is also why the free result and a connected result can reasonably differ once an account exists. A connected crawl surfaces more pages than the preview sample; paid market data replaces an inferred opportunity with demand and live-SERP evidence; and repeated ChatGPT and Gemini sampling replaces a single-point probe with a measured pattern over time. Site owners can separately compare Search Console data in their own account. The free snapshot's job is to be honest about which evidence it has, not to blur the two scopes together.

The preview never mutates the site

Preview data is tied to the account that started the scan. Connecting Shopify for supported page publishing, running deeper research, or approving execution happens after the snapshot.

That distinction — reading versus writing — is deliberate and absolute for the bounded scan itself: nothing in it submits a form, logs into an admin panel, or issues a request a site's own visitors couldn't also make by browsing it. The scan's only footprint is the same kind of request any browser or crawler already makes when it visits a public page.

Connecting a platform is a distinct, later, explicit step that requires its own authorization, and it's the point where the relationship with the site changes from reading what's public to seeing what's private — a Search Console property's real query data, a CMS's actual catalog — and, eventually, being able to propose or make changes. Nothing in the free snapshot implies that connection has already happened or that any change has been made on its behalf.

Common questions

How many pages does the free snapshot actually check?

Roughly 10–20 canonical public pages — the homepage and the pages a normal visitor or crawler would reach through primary navigation, not every URL, parameter variant, or paginated listing a larger site accumulates. It's a representative sample of the site's technical foundation, not a full-site audit.

Does the free snapshot see the same thing Google Search Console sees?

No. The snapshot checks observable, public conditions — titles, canonicals, robots and sitemap rules, structured data, response behavior — that are verifiable independently of any account. Actual indexing status, query data, and click behavior live in Search Console and require a connected, verified property; the free result can point at conditions that tend to cause indexing problems without being able to confirm which specific URLs are currently affected.

Will running the snapshot change anything on the site?

No. The bounded scan only reads what's already public — the same kind of request any browser or crawler already makes — and doesn't submit forms, sign in, or alter anything. Connecting Shopify is a separate, later, explicit step, and any proposed change happens after that, not during the free scan.

Sources

  1. Google Search Central — Introduction to robots.txt
  2. Google Search Central — Learn about sitemaps
  3. Google Search Central — What is canonicalization
  4. Google Search Central — How to specify a canonical URL with rel="canonical"
  5. Google Search Central — Influencing your title links in search results
  6. Google Search Central — General structured data guidelines
  7. Google Search Central — Intro to how structured data markup works
  8. Google Search Console Help — Page indexing report

Keep reading

Back to all guides