Research Methodology: SEO/AEO Experiment
Detailed explanation of the 2×2×2 factorial design, measurement protocol, observation schedule, and data interpretation approach.
Experiment Overview
This is a controlled research study testing how page characteristics affect search engine discovery, indexing, ranking, and AI retrieval in a low-competition market (Falkland Islands service enquiries).
Research Questions
- Does stronger internal linking reduce time to discovery, crawl, or indexing?
- Does richer contextual content improve ranking for natural-language queries after indexing?
- Does richer semantic HTML improve machine interpretation or eligibility for rich results?
- Do combinations of factors perform differently from individual factors?
- Does a page need to be indexed by a search engine before an AI system reliably retrieves or cites it?
- Do Google and Bing respond differently to the same controlled variants?
Experimental Design: 2×2×2 Factorial
We test three factors independently, each at two levels:
Factor H: Semantic HTML Richness
- Low (H=0):
- Valid HTML; unique title, meta description, H1, self-canonical, one main content region, ordinary headings and paragraphs; no page JSON-LD.
- High (H=1):
- All low-level controls plus article, header, nav, section, aside tags, breadcrumbs, definition/list/table/FAQ structures; valid WebPage or Article plus BreadcrumbList JSON-LD matching visible text.
Factor C: Contextual Depth
- Low (C=0):
- 350-500 words; answer-first summary, service definition, Falkland Islands relevance, selection cautions, concise FAQ.
- High (C=1):
- 900-1300 words; all low elements plus terminology variants, comparison with alternatives, local constraints, verification checklist, evidence/method notes, limitations, expanded FAQ.
Factor L: Internal Link Strength
- Low (L=0):
- Included in XML sitemap and linked once from its topic hub; no homepage or cross-content promotion.
- High (L=1):
- All low links plus descriptive links from experiment hub, FLK jurisdiction hub, one relevant service directory, one relevant editorial page, and reciprocal breadcrumb/topic navigation.
Total combinations per topic: 2 × 2 × 2 = 8
Total experimental pages: 8 (Tui Na) + 8 (Vietnamese massage) = 16
Content Requirements
Truthfulness and Safety
- No page invents a provider, street address, phone number, review, rating, credential, opening hour, price, or testimonial.
- No page diagnoses a condition, prescribes treatment, promises pain relief or recovery, or implies massage replaces qualified medical care.
- Health-related language is framed as general information and service-selection guidance.
- Red-flag symptoms direct readers to appropriately qualified healthcare professionals.
Research Disclosure
Every page includes:
- Unique research token (e.g., LCI-EXP-T-000) for retrieval measurement
- Statement that the page is part of a controlled research experiment
- Disclosure that it is not a real provider listing
- Page publication date and last-reviewed date
- Link to this methodology page
- Safety notice about health information
Uniqueness
Every page has a unique scenario, title, meta description, H1, opening answer, examples, and FAQ questions. No paragraphs are copied unchanged between variants (except standardized disclosures).
Measurement Protocol
Observation Schedule
| Checkpoint | Days After Publication | Observations |
|---|---|---|
| D0 (Publication) | 0 | HTTP status, canonical, robots directive, sitemap presence |
| D1 | 1 | Crawl evidence, Google URL Inspection, Bing URL Inspection |
| D3 | 3 | Index status, keyword queries, rank checks |
| D7 | 7 | Index rate, natural-query visibility, AI retrieval trials |
| D14 | 14 | Index rate, rank progression, AI citation checks |
| D28 | 28 | Final index rate, visibility summary, citation analysis |
Important: If a page first becomes indexed after D28, we record the event and begin a separate 14-day post-index AI observation window. Pre-index AI trials are separated into a clearly labelled exploratory dataset.
Query Classes
Each page has stable query IDs in four classes:
- Q-TOKEN:
- Exact research token (e.g., "LCI-EXP-T-000") for unambiguous page detection
- Q-TITLE:
- Exact page title for index/retrieval confirmation
- Q-NEAR:
- Close topical wording without research token
- Q-NATURAL:
- Realistic conversational service discovery (e.g., "How could I look for Chinese Tui Na in the Falkland Islands?")
Search Engine Protocol
- Test Google and Bing separately
- Record desktop/mobile mode
- Record signed-in/incognito state and approximate tester location
- For rank, record exact URL and position; if position unclear, use fixed ranges: 1-10, 11-20, 21-50, 51-100, not found
- Save snippets when they demonstrate correct or incorrect page interpretation
AI Retrieval Protocol
- Test selected versions of ChatGPT, Gemini, DeepSeek, and Bing/Copilot
- Use a new conversation for every trial
- Record whether live web search/browsing is enabled
- Run three independent trials per page/query/model/checkpoint when practical
- Score each trial: 0=absent, 1=domain mentioned only, 2=page-specific retrieval without citation, 3=exact canonical citation
- Save complete prompt, answer, citations, model label, date, and settings
Note: AI testing begins only after confirmed page indexing. Pre-index trials are exploratory only.
Definitions and Outcome Funnel
| Stage | Definition | Evidence Required |
|---|---|---|
| Accessible | Production URL returns HTTP 200 with usable HTML | Timestamped status check and page render |
| Discoverable | URL in approved sitemap or reachable from allowed entry point | Sitemap record and link graph |
| Submitted | URL submitted via Search Console, Webmaster Tools, or IndexNow | Submission timestamp, method, response |
| Crawled | Search engine reports or logs a crawl event | Search Console inspection or verified crawler log |
| Indexed | Engine explicitly reports indexed status or returns URL for controlled query | Inspection evidence preferred; site: queries secondary |
| Ranked | Exact experimental URL appears for predefined query | Position/range, evidence link or screenshot |
| Retrieved | Search-enabled model uses page-specific facts in its answer | Saved prompt and answer with evidence |
| Cited | Model presents clickable citation to exact experimental URL | Exact citation URL and saved answer |
Success Thresholds (by D28)
- Minimum Viable:
- At least 12 of 16 pages indexed by Google or Bing; all pages technically valid; at least one non-brand impression or visible result per topic
- Useful:
- At least 12 pages indexed, at least 6 visible for natural/near query, and at least 2 exact page citations in post-index AI trials
- Strong:
- At least 14 pages indexed; high-C or high-L variants show consistent directional advantage in both topics; citations occur on at least two model families
Important: Failure to meet a threshold is still a valid result if implementation and observation criteria are met. We report evidence, not claims of certainty beyond what data supports.
Known Limitations
- Small sample: Two topic blocks only. Results are directional evidence, not statistically robust population estimates.
- Low competition: Falkland Islands is a very small market. Results may not generalize to higher-competition locations.
- Synthetic content: Pages are research pages, not real business operations. Real provider visibility may differ.
- 28-day window: Search engines may continue discovering and indexing beyond this period. We do not claim this is a definitive result.
- Model variability: AI systems change versions frequently. Results may not be replicable with future models.
- Tester location and settings: Observed results depend on tester location, signed-in state, safe-search settings, and search history.
AI Assistance Disclosure
Generative AI (Claude 3.5 Sonnet) assisted in drafting experimental page content. Each page was reviewed by the research team for factual accuracy, safety, uniqueness, and compliance with research requirements before publication. AI assistance does not constitute endorsement of the information as medical advice or professional guidance.
Questions or Feedback
If you have questions about this methodology or want to provide feedback, please contact us.