πŸ₯ HIST

Health Information Source Trustworthiness – About the tool

Version 3.5 β€” Normalized Scoring with Strengths & Email Dashboard

πŸ“‹ What is this tool?

HIST (Health Information Source Trustworthiness) is an automated trustworthiness checker for health and medical websites. It analyzes technical signals that indicate credibility – not medical accuracy.

🎯 Core mission: Help users identify websites that follow trust and transparency standards (authorship, citations, dates, disclaimers, structured data).

The tool does NOT evaluate medical correctness – it evaluates trust signals that correlate with reliable health information.

βœ… What we check (technical & trust signals)

CRITICAL (20pts) HIGH (15pts) MEDIUM (10pts) LOW (5pts) INFO (0pts)

πŸ“Š Scoring Formula β€” Version 2.0

πŸ”„ What's new: The scoring system now uses normalized scoring – your score reflects the percentage of trust signals your page fulfills, with each signal weighted by importance. This creates a finer gradation from 0 to 100 and avoids penalizing pages too harshly.

How it works

score = (1 βˆ’ totalDeduction / MAX_DEDUCTION) Γ— 100 + historicalBonus

Where:

Weight distribution

Category Rules Deduction per rule Total possible deduction
CRITICAL 5 (medicalwebpage, medicalentity, author, review, citations) 20 pts 100 pts
HIGH 3 (jsonld, specialty, audience) 15 pts 45 pts
MEDIUM 5 (publisher, disclaimer, dates, structure_h1, structure_h2) 10 pts 50 pts
LOW 6 (author_credentials, citations_weak, dates_modified, structure_paragraphs, semantic_html, health_relevance) 5 pts 30 pts
INFO 5 (structure_lists, internal_links, llms_txt, markdown, provenance) 0 pts 0 pts
TOTAL 225 pts

Score interpretation

90-100: Excellent trustworthiness 70-89: Good – room for improvement 50-69: Moderate – verify references 0-49: Low trustworthiness – do not rely

πŸ“œ Historical document bonus

+15 bonus points – Older but landmark medical papers (e.g., classic clinical trials, seminal research) may receive a bonus because age alone does not reduce trustworthiness for historically significant work. The bonus is applied after normalization but capped at 100.

πŸ“ Example calculations

Scenario Missing rules totalDeduction Score (Ver. 2.0) Score (Ver. 1.5)
3 critical rules missing medicalwebpage (-20), author (-20), jsonld (-15) 55 76 45
5 critical rules missing medicalwebpage (-20), medicalentity (-20), author (-20), review (-15), citations (-15) 90 60 10
2 critical + 3 medium medicalwebpage (-20), author (-20), publisher (-10), disclaimer (-10), dates (-10) 70 69 30
All rules fulfilled None 0 100 100
All rules missing All 19 rules 225 0 0
Landmark paper (old + bonus) Missing JSON-LD (-15), missing dates (-10) 25 + bonus 15 89 90

⚠️ Important limitations – PLEASE READ

This tool has fundamental limitations that every user must understand:

πŸ”’ 1. Protected websites (cannot be analyzed automatically)

The following websites block automated requests. For these, the tool provides manual verification checklists instead of automatic scoring:

β†’ For these sites, we provide trust estimates based on domain reputation and step‑by‑step manual verification instructions.

🧠 2. We do NOT check medical accuracy

No automated tool can verify if medical information is factually correct. Our tool checks trust signals (authorship, citations, dates, disclaimers), not whether the content is medically accurate.

Example: A well‑structured page with an MD author, citations, and JSON-LD could still contain incorrect medical advice. Always consult a real doctor.

πŸ” 3. We cannot verify credentials

If a page claims "Dr. John Smith, MD" – we detect the title, but we cannot verify that John Smith actually holds an MD degree. Anyone can add "MD" to their name online.

πŸ“± 4. JavaScript‑heavy websites

Modern React/Vue/Angular sites that render content dynamically may not be fully analyzable. We only see the initial HTML response.

πŸ’Έ 5. Paywalled content

If a page requires login or payment, we can only analyze metadata in the <head> section, not the full content.

πŸ“… 6. Age vs. trustworthiness

Older content is not automatically untrustworthy. Classical medical literature and landmark clinical trials remain highly valuable. The tool applies a historical document bonus for significant older papers.

πŸ“š Standards & sources

Our methodology is based on established trust frameworks:

πŸ“– References:
β€’ HONcode: https://www.honcode.com/
β€’ Google E-E-A-T: Google Search Central
β€’ Schema.org MedicalWebPage: https://schema.org/MedicalWebPage

πŸ’‘ When to use this tool

βœ… Small health blogs βœ… New medical websites βœ… SEO optimization for health content βœ… Quick credibility check for unfamiliar sites

🚫 When NOT to rely on this tool

❌ Making medical decisions ❌ Replacing professional medical advice ❌ Verifying factual accuracy of treatments ❌ Analyzing paywalled or protected sites
⚠️ Final disclaimer: This tool is for educational and informational purposes only. Always consult a qualified healthcare professional for medical decisions.

πŸ“Š Research & Data Usage

πŸ”¬ Important notice on data collection and research use:

By using this tool, you acknowledge that anonymized scan results (URLs, scores, detected issues, and aggregated statistics) may be used for scientific research purposes. This research aims to:

  • Better understand current practices in health information publishing,
  • Identify patterns of trustworthiness across different types of health websites,
  • Develop improved guidelines and standards for health information sources,
  • Contribute to the development of more reliable AI-based health information systems.

No personal data (email addresses) will ever be shared or published – only aggregated, anonymized statistics. You may request deletion of your email-associated data at any time by contacting us.

Data retention: All scan records are permanently stored in our secure database to enable longitudinal research on health information quality trends.