πŸ“š Documentation & References

Complete reference guide to all trust signals, their sources, and implementation standards

πŸ“‹ Table of Contents

πŸ“Š Scoring Methodology β€” Version 2.0

πŸ”„ What's new: The scoring system now uses normalized scoring – your score reflects the percentage of trust signals your page fulfills, with each signal weighted by importance. This creates a finer gradation from 0 to 100 and avoids penalizing pages too harshly.

Formula

score = (1 βˆ’ totalDeduction / MAX_DEDUCTION) Γ— 100 + historicalBonus

Where:

Weight distribution

Category Rules Deduction per rule Total possible deduction
CRITICAL 5 (medicalwebpage, medicalentity, author, review, citations) 20 pts 100 pts
HIGH 3 (jsonld, specialty, audience) 15 pts 45 pts
MEDIUM 5 (publisher, disclaimer, dates, structure_h1, structure_h2) 10 pts 50 pts
LOW 6 (author_credentials, citations_weak, dates_modified, structure_paragraphs, semantic_html, health_relevance) 5 pts 30 pts
INFO 5 (structure_lists, internal_links, llms_txt, markdown, provenance) 0 pts 0 pts
TOTAL 225 pts

Score interpretation

90-100
Excellent trustworthiness β€” Highly recommended
70-89
Good β€” Room for improvement
50-69
Moderate β€” Verify references before use
0-49
Low trustworthiness β€” Do not rely on this source

πŸ“œ Historical document bonus

+15 bonus points – Older but landmark medical papers (e.g., classic clinical trials, seminal research) may receive a bonus because age alone does not reduce trustworthiness for historically significant work. The bonus is applied after normalization but capped at 100.
πŸ“– Source: Landmark Clinical Trials β€” Many foundational medical studies are decades old but remain highly cited and trusted.

πŸ‘₯ Human Credibility Signals

These signals are designed to help human readers assess trustworthiness through authorship, citations, transparency, and readability.

πŸ“ Author with Credentials
Human Signal CRITICAL -20 pts

What we check: Presence of author name with medical credentials (MD, PhD, specialist certification) and affiliation.

Why it matters: Readers need to know who wrote the content. 78% of users check authorship before trusting medical information online. Anonymous content is 3x more likely to be considered unreliable.

πŸ“– Source: Google E-E-A-T Guidelines β€” Expertise is one of the four pillars of trust assessment.
πŸ“– Source: HONcode Principle #1: Authority β€” Medical advice should only be given by qualified professionals.
βœ… Example: "Dr. Ana PetroviΔ‡, MD, Endocrinologist at University Hospital Belgrade"
πŸŽ“ Author Credentials (Titles)
Human Signal LOW -5 pts

What we check: Author name includes medical credentials (MD, PhD, Board Certified, etc.).

Why it matters: Credentials provide evidence of expertise. Google's E-E-A-T guidelines specifically require "Expertise" signals for YMYL (Your Money or Your Life) content.

πŸ“– Source: Google Search Quality Rater Guidelines β€” YMYL content requires high levels of E-E-A-T.
βœ… Good: "Dr. Ana PetroviΔ‡, MD, PhD, FRCP" β€” provides clear expertise signals
❌ Bad: "Ana PetroviΔ‡" β€” no credentials
πŸ”¬ Medical Reviewer
Human Signal HIGH -15 pts

What we check: Presence of a named medical reviewer with credentials (reviewedBy in JSON-LD or visible text).

Why it matters: Medical review adds accountability. 65% of users trust content more when they see it was reviewed by a medical professional.

πŸ“– Source: HONcode Principle #2: Complementarity β€” Information should support but not replace the doctor-patient relationship.
βœ… Good: "Medically reviewed by Dr. Marko JovanoviΔ‡, MD, PhD β€” Last reviewed: March 2026"
πŸ“„ Citations & References
Human Signal HIGH -15 pts

What we check: Links to PubMed, WHO, government (.gov) sources, or DOIs (Digital Object Identifiers).

Why it matters: Citations provide verifiable evidence. 82% of readers check sources before trusting medical claims. Missing citations is the #1 red flag for medical misinformation.

πŸ“– Source: HONcode Principle #4: Attribution β€” Sources and dates must be clearly stated.
πŸ“– Source: WHO Fact Sheet Standards β€” All WHO health information includes cited sources.
βœ… Good: "According to a 2023 study published in The Lancet (DOI: 10.1016/...)"
❌ Bad: "According to experts..." β€” no verifiable source
πŸ“„ Weak References
Human Signal LOW -5 pts

What we check: Citations exist but are not from trusted sources (PubMed, WHO, .gov, DOI).

Why it matters: Not all sources are equal. Blog posts and news articles are not considered authoritative medical sources.

πŸ“– Source: WHO Publishing Standards β€” Only peer-reviewed or official sources are used in WHO publications.
❌ Weak: "According to a blog post..."
βœ… Strong: "According to a 2024 randomized controlled trial published in NEJM (DOI: 10.1056/...)"
πŸ›οΈ Publisher as MedicalOrganization
Human Signal MEDIUM -10 pts

What we check: Publisher is clearly identified as a medical organization (hospital, university, research institute).

Why it matters: Knowing the publisher helps readers assess institutional authority. 72% of users trust medical info more when it comes from a known hospital or research institution.

πŸ“– Source: HONcode Principle #6: Transparency of authorship β€” Clear identification of the organization publishing the content.
βœ… Good: "Published by Mayo Clinic, a nonprofit academic medical center."
βš–οΈ Medical Disclaimer
Human Signal MEDIUM -10 pts

What we check: Clear disclaimer stating content is informational and not a substitute for professional medical advice.

Why it matters: Legal protection and user safety. 58% of users may mistakenly use online medical info as professional advice. A clear disclaimer is both responsible and required by law in many jurisdictions.

πŸ“– Source: HONcode Principle #8: Complementarity β€” The information provided should support, not replace, the doctor-patient relationship.
πŸ“– Source: FDA Medical Device Guidance β€” Medical information must include appropriate disclaimers.
βœ… Good: "This information is not intended to replace professional medical advice. Always seek the advice of your physician."
πŸ“… Publication Date
Human Signal MEDIUM -10 pts

What we check: Published date (datePublished in JSON-LD or visible text).

Why it matters: Medical knowledge evolves rapidly. 89% of users check the date before reading medical content. Articles older than 3 years may be outdated.

πŸ“– Source: HONcode Principle #4: Attribution β€” Dates of all content must be clearly stated.
βœ… Good: "Published: February 10, 2025"
❌ Bad: No publication date visible
πŸ”„ Last Modified Date
Human Signal LOW -5 pts

What we check: Last modified/updated date (dateModified in JSON-LD or visible text).

Why it matters: Modified dates show content is actively maintained. Outdated medical info can be dangerous. 74% of users prefer sites that show when content was last reviewed.

πŸ“– Source: WHO Publishing Standards β€” Guidelines for regular content review and update cycles.
βœ… Good: "Last updated: March 1, 2026 β€” Reviewed by Dr. Smith"
πŸ“‘ H1 Heading Structure
Human Signal LOW -5 pts

What we check: Presence of <h1> heading on the page.

Why it matters: H1 headings help users and search engines understand the page's main topic. 65% of users scan headings before reading. A missing H1 reduces accessibility and SEO.

πŸ“– Source: W3C Accessibility Guidelines β€” Headings are essential for screen readers and accessibility.
βœ… Good: <h1>Understanding Type 2 Diabetes</h1>
πŸ“‘ H2 Subheadings
Human Signal LOW -5 pts

What we check: Presence of <h2> subheadings organizing content.

Why it matters: H2 headings organize content into scannable sections. 79% of users scan rather than read entire articles. Clear H2 headings make your content more accessible.

πŸ“– Source: Nielsen Norman Group β€” Users scan content in an F-shaped pattern; headings help guide this scan.
βœ… Good: <h2>Common Symptoms</h2><h2>Treatment Options</h2>
πŸ“‘ Paragraph Length
Human Signal LOW -5 pts

What we check: Paragraphs over 150 words in length.

Why it matters: Long paragraphs are hard to read on screens. 55% of users will skip over paragraph blocks of text. Shorter paragraphs improve readability and engagement.

πŸ“– Source: Nielsen Norman Group β€” Web users read 20-28% of words on average; shorter paragraphs increase comprehension.
❌ Bad: 200-word wall of text.
βœ… Good: Multiple short paragraphs of 2-4 sentences each.
πŸ“‘ Lists & Bullet Points
Human Signal INFO 0 pts

What we check: Presence of <ul> or <ol> lists.

Why it matters: Lists make content scannable. 70% of users prefer content with lists and bullet points. They also help with AI/LLM understanding of key points.

πŸ“– Source: Plain Language Guidelines β€” Lists are recommended for clear communication of complex information.
βœ… Good: <ul><li>Increased thirst</li><li>Frequent urination</li></ul>
πŸ₯ Health Relevance
Human Signal INFO 0 pts

What we check: Presence of medical keywords in content (diabetes, cancer, treatment, symptoms, etc.).

Why it matters: Medical keywords help users find your content and help AI understand the topic. They also signal that your content is health-focused.

πŸ“– Source: MedlinePlus Health Literacy β€” Using clear medical terminology improves patient understanding.
βœ… Good: Using terms like "diabetes", "metformin", "clinical trial", "glycemic control" naturally in text.

πŸ€– LLM & AI Readiness Signals

These signals help AI systems and LLMs properly understand, categorize, and trust your medical content.

πŸ€– MedicalWebPage Type
Machine Signal CRITICAL -20 pts

What we check: @type: MedicalWebPage in JSON-LD structured data.

Why it matters: AI systems need explicit signals. MedicalWebPage tells AI this is medically relevant content, not just general articles. Without this, AI may treat it as generic content.

πŸ“– Source: Schema.org MedicalWebPage β€” Standardized definition for medical web pages.
πŸ“– Source: Google Structured Data Guide β€” How structured data helps search engines understand your content.
βœ… Good: {"@type":"MedicalWebPage","medicineSystem":"Evidence-based medicine","medicalAudience":"Patient"}
πŸ€– Medical Entity Definition
Machine Signal CRITICAL -20 pts

What we check: Medical entity type defined (MedicalCondition, MedicalTherapy, MedicalGuideline, ScholarlyArticle).

Why it matters: AI systems need to know the specific medical subject. Defining the entity helps AI correctly categorize and understand the content.

πŸ“– Source: Schema.org MedicalEntity β€” Parent type for all medical concepts in Schema.org.
πŸ“– Source: Schema.org Health & Lifesciences β€” Complete medical vocabulary for structured data.
βœ… Good: {"@type":"MedicalCondition","name":"Type 2 Diabetes Mellitus"}
πŸ€– JSON-LD Structured Data
Machine Signal HIGH -15 pts

What we check: Presence of <script type="application/ld+json"> tags.

Why it matters: JSON-LD is the standard for machine-readable data. AI systems like Google's search use it to understand page context. Without it, AI has to guess.

πŸ“– Source: JSON-LD Specification β€” W3C standard for linked data in JSON.
πŸ“– Source: Google Structured Data β€” How JSON-LD is used in search results.
βœ… Good: <script type="application/ld+json">{"@context":"https://schema.org","@type":"MedicalWebPage"}</script>
πŸ€– Medical Specialty
Machine Signal MEDIUM -10 pts

What we check: relevantSpecialty property in JSON-LD.

Why it matters: MedicalSpecialty helps AI understand the domain. A cardiology page and an endocrinology page should be treated differently. It improves AI accuracy.

πŸ“– Source: Schema.org MedicalSpecialty β€” Standardized medical specialty vocabulary.
βœ… Good: {"relevantSpecialty":{"@type":"MedicalSpecialty","name":"Cardiology"}}
πŸ€– Medical Audience
Machine Signal MEDIUM -10 pts

What we check: medicalAudience property in JSON-LD.

Why it matters: MedicalAudience tells AI who the content is for. Patient content should be simpler, physician content more technical. This improves AI recommendations.

πŸ“– Source: Schema.org MedicalAudience β€” Standardized audience type for medical content.
βœ… Good: {"medicalAudience":"Patient"} β€” for public health content
πŸ€– Semantic HTML
Machine Signal LOW -5 pts

What we check: Use of semantic HTML elements (<article>, <section>, <nav>, <header>, <footer>, <main>, <aside>).

Why it matters: Semantic HTML helps AI understand page structure. <article> clearly indicates main content. <section> organizes sections. This improves accessibility and AI understanding.

πŸ“– Source: W3C HTML5 Semantics β€” Official HTML5 semantic element definitions.
βœ… Good: <article><header>Title</header><section>Content</section></article>
πŸ€– Internal Linking
Machine Signal INFO 0 pts

What we check: Number of internal links (links to other pages on the same site).

Why it matters: Internal links create a knowledge graph. AI can understand relationships between topics. Users can navigate to related content. It also helps with SEO.

πŸ“– Source: Google SEO Starter Guide β€” Internal linking helps search engines understand site structure.
βœ… Good: <a href="/diabetes-type-1">Learn about Type 1 Diabetes</a>
πŸ€– llms.txt File
Machine Signal INFO 0 pts

What we check: Presence of link to /llms.txt file.

Why it matters: /llms.txt is a standard AI entry point. It tells AI which pages are most important. Think of it as a sitemap specifically for AI systems.

πŸ“– Source: llms.txt Standard β€” Community-driven standard for AI site maps.
βœ… Good: <a href="/llms.txt">llms.txt for AI</a>
πŸ€– Markdown Version
Machine Signal INFO 0 pts

What we check: Link to .md markdown version of the content.

Why it matters: AI systems parse markdown more efficiently than HTML. Clean markdown reduces parsing errors and helps AI extract key information.

πŸ“– Source: Markdown Standard β€” Original Markdown specification.
βœ… Good: <a href="/content.md">Download Markdown version</a>
πŸ€– PROV-O Provenance
Machine Signal INFO 0 pts

What we check: PROV-O structured data (prov:Activity type).

Why it matters: PROV-O provides detailed provenance (who did what, when). AI systems can track content history. This is particularly important for medical content that gets updated regularly.

πŸ“– Source: W3C PROV-O Specification β€” Official W3C provenance ontology.
βœ… Good: {"@context":"http://www.w3.org/ns/prov","@type":"Activity","wasAssociatedWith":{"name":"Dr. Smith"}}

πŸ“š Standards & Frameworks

Our methodology is based on established trust frameworks:

1. HONcode (Health on the Net)

Eight principles for trustworthy health websites:

  1. Authority: Medical advice only by qualified professionals
  2. Complementarity: Information supports, doesn't replace, doctor-patient relationship
  3. Privacy: Confidentiality of patient data
  4. Attribution: Sources and dates clearly stated
  5. Justification: Claims must be backed by evidence
  6. Transparency of authorship: Clear identification of publisher
  7. Transparency of funding: Sponsors and advertisers clearly identified
  8. Advertising: Clear distinction between editorial and advertising content
πŸ“– Source: HONcode Official Website

2. Google E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness)

Four pillars for assessing content quality, especially for YMYL (Your Money or Your Life) topics like health:

  • Experience: First-hand or life experience with the topic
  • Expertise: Formal knowledge, skills, and qualifications
  • Authoritativeness: Reputation and recognition in the field
  • Trustworthiness: Accuracy, honesty, and safety of content

3. Schema.org Medical Vocabulary

Standardized structured data for medical content:

4. W3C PROV-O (Provenance Ontology)

Standard for tracking provenance and content history:

  • Activity: Actions or events
  • Agent: Who performed the activity
  • Entity: What was created or modified
πŸ“– Source: W3C PROV-O Specification

5. Additional Standards

⚠️ Limitations

Please read carefully: This tool has fundamental limitations that every user must understand.

πŸ”¬ We do NOT check medical accuracy

No automated tool can verify if medical information is factually correct. Our tool checks trust signals (authorship, citations, dates, disclaimers), not whether the content is medically accurate.

Example: A well-structured page with an MD author, citations, and JSON-LD could still contain incorrect medical advice. Always consult a real doctor.

πŸ”’ Protected websites

The following websites block automated requests:

  • PubMed / NCBI / NIH β€” Returns "Checking your browser" protection
  • Elsevier / ScienceDirect / Springer / Nature β€” Academic paywalls and bot protection
  • Any site with Cloudflare, DDoS guard, or CAPTCHA

β†’ For these sites, we provide trust estimates based on domain reputation and step-by-step manual verification instructions.

🧠 We cannot verify credentials

If a page claims "Dr. John Smith, MD" – we detect the title, but we cannot verify that John Smith actually holds an MD degree. Anyone can add "MD" to their name online.

πŸ“± JavaScript-heavy websites

Modern React/Vue/Angular sites that render content dynamically may not be fully analyzable. We only see the initial HTML response.

πŸ’Έ Paywalled content

If a page requires login or payment, we can only analyze metadata in the <head> section, not the full content.

πŸ“… Age vs. trustworthiness

Older content is not automatically untrustworthy. Classical medical literature and landmark clinical trials remain highly valuable. The tool applies a historical document bonus for significant older papers.

⚠️ Final disclaimer: This tool is for educational and informational purposes only. Always consult a qualified healthcare professional for medical decisions.