ToolUniverse Disease Research
Generate a comprehensive disease research report with full source citations. The report is created as a markdown file and progressively updated during research.
IMPORTANT: Always use English disease names and search terms in tool calls. Respond in the user's language.
LOOK UP, DON'T GUESS
When asked about a disease, query Orphanet/OMIM/DisGeNET FIRST. Don't rely on memory for prevalence, genetics, or treatment — these change over time. When you're not sure about a fact, your first instinct should be to SEARCH for it using tools, not to reason harder from memory.
When to Use
- User asks about any disease, syndrome, or medical condition
- Needs comprehensive disease intelligence or a detailed research report
- Asks "what do we know about [disease]?"
Core Workflow: Report-First Approach
DO NOT show the search process to the user. Instead:
- Create report file first - Initialize
{disease_name}_research_report.md - Research each dimension - Use all relevant tools
- Update report progressively - Write findings after each dimension
- Include citations - Every fact must reference its source tool
Disease Mechanism Reasoning
When synthesizing disease etiology, trace the full pathogenic cascade:
- Genetic basis - Which variants (rare or common) confer risk, and in which genes?
- Molecular mechanism - How do those variants alter protein function, expression, or regulation?
- Cellular effect - What downstream cellular processes are disrupted (signaling, metabolism, stress response)?
- Tissue/organ manifestation - How does cellular dysfunction present as organ-level pathology?
This chain structures the Genetic & Molecular Basis (Section 3) and Biological Pathways (Section 5) sections.
10 Research Dimensions
See: tool_usage_details.md for complete tool calls per section.
Normalizing free text to ontology IDs (Dimension 1)
When the input is messy free text (a sample attribute, a synonym, a tissue/organism label) rather than a clean disease name, use ZOOMA_annotate_text to map it to standardized ontology terms (EFO/MONDO/UBERON/etc.) before lookup. It returns each match as an ontology IRI with a confidence rating (HIGH/GOOD/MEDIUM/LOW), so you can keep only high-confidence hits and feed the resolved ID into OLS / OpenTargets.
Each match also carries a ready-to-use curies field (e.g. MONDO:0004979) so you can feed the resolved ID straight into OLS / OpenTargets without parsing the IRI. ZOOMA is the live replacement for the retired OxO cross-reference service; pair it with ols_get_efo_term to expand the resolved IRI into labels, synonyms, and hierarchy.
Report Template
Create this file structure at the start:
Citation Format
Every piece of data MUST include its source:
In tables: Add a Source column with tool name
In lists: - Finding [Source: tool_name]
In prose: (Source: tool_name, query: "...")
References section: Complete tool usage log with parameters
Progressive Update Pattern
Evidence Grading & Interpretation
Every finding in the report should be graded:
Synthesis Questions (answer in Executive Summary)
After collecting data from all 10 dimensions, the report MUST answer:
- What causes this disease? Summarize the genetic architecture (monogenic vs polygenic, key loci, penetrance)
- What are the therapeutic options? Ranked by evidence level and approval status
- What biomarkers exist? For diagnosis, prognosis, and treatment selection
- What's the unmet need? What aspects lack effective treatment or understanding?
- What are the active research frontiers? Based on clinical trials and recent publications
Interpreting Cross-Database Concordance
When multiple databases provide different data for the same disease:
- OpenTargets + DisGeNET + OMIM agree on a gene: T1 evidence — high confidence
- Only OpenTargets reports an association: Check the datasource scores — genetic_association > literature > animal_model
- DisGeNET score > 0.5 but not in OpenTargets: May be text-mined; verify with PubMed
- Gene in GWAS but not OMIM: Likely a complex disease susceptibility locus, not Mendelian
Handling Conflicting Data
Final Report Quality Checklist
- All 10 sections have content (or marked "No data available")
- Every data point has a source citation
- Executive summary reflects key findings
- References section lists all tools used
- Tables properly formatted
- No placeholder text remains
Expected Output Scale
For a well-studied disease (e.g., Alzheimer's), the final report should include:
- 5+ ontology IDs, 10+ synonyms, disease hierarchy
- 20+ phenotypes with HPO IDs
- 50+ genes, 30+ GWAS associations, 100+ ClinVar variants
- 20+ drugs, 50+ clinical trials
- 10+ pathways, PPI network, expression data
- 100+ publications
- 15+ similar diseases
- Drug warnings and adverse events
Total: 500+ individual data points, each with source citation.
Cross-Skill References
For rare disease differential diagnosis, run: python3 skills/tooluniverse-rare-disease-diagnosis/scripts/clinical_patterns.py --type differential --symptoms 'symptom1,symptom2'
Reference Files
- REPORT_TEMPLATE.md [blocked] - Full report markdown template and citation format guide
- RESEARCH_PROTOCOL.md [blocked] - Step-by-step code procedures, progressive update pattern, quality checklist
- tool_usage_details.md [blocked] - Complete tool calls for each research dimension
- TOOLS_REFERENCE.md [blocked] - Complete tool documentation
- EXAMPLES.md [blocked] - Sample disease research reports

