PubMed Database
Credit: This skill comes from a community author. Keep this credit when you copy or change it.
Use this skill when the source must be PubMed. Do not use a broad web search in place of PubMed.
Use This Skill For
- Finding health and life science papers.
- Building searches with MeSH terms and field tags.
- Finding PMIDs, abstracts, authors, dates, and journal data.
- Finding linked or similar papers.
- Making a search that another person can repeat.
- Using NCBI E-utilities from Python, a shell, or an HTTP client.
Build the Search
Split the question into key ideas. Add terms for each idea. Join them with uppercase Boolean words.
(concept_1 OR synonym_1) AND (concept_2 OR synonym_2)
concept_1 AND filter
concept_1 NOT excluded_termUse parentheses when you mix AND and OR. Without them, PubMed may read the search in a way you did not mean.
Useful field tags:
[ti]: title[ab]: abstract[tiab]: title or abstract[au]: author[ta]: journal short name[mh]: MeSH term[majr]: main MeSH topic[pt]: paper type[dp]: publication date[la]: language
Examples:
diabetes mellitus[mh] AND treatment[tiab] AND systematic review[pt] AND 2023:2026[dp]
(metformin[nm] OR insulin[nm]) AND diabetes mellitus, type 2[mh] AND randomized controlled trial[pt]
smith ja[au] AND cancer[tiab] AND 2026[dp] AND english[la]Check the final search in PubMed. Read the Search Details box. It shows how PubMed mapped each term.
Use MeSH With Care
Use MeSH when the idea has a clear MeSH term. For a new topic, use both MeSH and plain words.
("Diabetes Mellitus, Type 2"[mh] OR "type 2 diabetes"[tiab] OR T2D[tiab])
AND
("Artificial Intelligence"[mh] OR "artificial intelligence"[tiab] OR "machine learning"[tiab])Use this form for a MeSH subheading:
diabetes mellitus, type 2/drug therapy[mh]
cardiovascular diseases/prevention & control[mh]Use [majr] only when the idea must be a main topic. It can remove useful papers.
MeSH is added after a paper is indexed. Very new papers may not have MeSH terms yet. Plain title and abstract terms help find them.
Add Filters
Paper type:
clinical trial[pt]
meta-analysis[pt]
randomized controlled trial[pt]
review[pt]
systematic review[pt]
guideline[pt]Publication date:
2026[dp]
2020:2026[dp]
2026/03/15[dp]Other filters:
english[la]
free full text[sb]
hasabstract[text]Do not add a free full text filter unless the task needs it. That filter can hide useful papers.
A publication type may not be added to a new paper at once. For a full search, add plain words too:
(randomized controlled trial[pt] OR random*[tiab])Use E-utilities
The main E-utilities steps are:
esearch.fcgifinds PMIDs.esummary.fcgigets short record data.efetch.fcgigets abstracts or full records.elink.fcgifinds linked and related records.
Use db=pubmed. Set retmode and rettype on purpose. Do not assume the default output will stay the same.
For a large result set, use usehistory=y. Pass WebEnv and query_key to later calls. Fetch results in small groups with retstart and retmax.
Concrete Python Example
This example searches PubMed and returns up to 20 PMIDs.
import os
import time
import requests
BASE = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
TOOL = "pubmed-database-skill"
def search_pubmed(query: str, retmax: int = 20) -> list[str]:
if not query.strip():
raise ValueError("The PubMed query is empty.")
params = {
"db": "pubmed",
"term": query,
"retmode": "json",
"retmax": max(1, min(retmax, 10_000)),
"tool": TOOL,
}
email = os.environ.get("NCBI_EMAIL")
if email:
params["email"] = email
api_key = os.environ.get("NCBI_API_KEY")
if api_key:
params["api_key"] = api_key
for attempt in range(3):
response = requests.get(
f"{BASE}/esearch.fcgi",
params=params,
timeout=30,
)
if response.status_code == 429 or response.status_code >= 500:
if attempt == 2:
response.raise_for_status()
time.sleep(2 ** attempt)
continue
response.raise_for_status()
data = response.json()
result = data.get("esearchresult", {})
return result.get("idlist", [])
return []
query = (
"hypertension[mh] "
"AND randomized controlled trial[pt] "
"AND 2024:2026[dp]"
)
pmids = search_pubmed(query)
print(pmids)Store the API key in NCBI_API_KEY. Store the contact email in NCBI_EMAIL. Never put a key in code, a saved command, a log, or a Git commit.
Follow NCBI rate limits. Space calls apart. Retry 429 and short server errors with a longer wait each time. Stop after a set number of tries.
Handle these cases:
- No results.
- A missing abstract.
- A paper with more than one date.
- A retracted paper or a correction.
- More results than
retmax. - Bad JSON or XML.
- A timeout.
- A
429rate-limit reply. - A
500class server reply. - Duplicate PMIDs from joined search paths.
Do not treat a PMID as proof that a paper is sound. Check the paper type, date, retraction state, and study details.
Example Task
Task: Find recent reviews about CRISPR treatment for sickle cell disease.
Search:
(
"Anemia, Sickle Cell"[mh]
OR "sickle cell disease"[tiab]
)
AND
(
"CRISPR-Cas Systems"[mh]
OR CRISPR[tiab]
OR "gene editing"[tiab]
)
AND
(
systematic review[pt]
OR review[pt]
OR review[ti]
)
AND
2020:2026[dp]Then:
- Run the search in PubMed.
- Check Search Details for bad term mapping.
- Save the exact search and result count.
- Fetch the PMIDs and record data.
- Remove duplicate PMIDs.
- Mark retractions and corrections.
- Record any paper removed by hand.
Keep a Search Log
For each search path, save:
- Exact search text.
- Database name.
- Search date.
- Filters.
- Result count.
- Export format.
- Any hand-made changes.
- Any papers removed by hand, with a reason.
Example:
| Database | Search date | Query | Filters | Results |
| --- | --- | --- | --- | ---: |
| PubMed | 2026-05-11 | `sickle cell disease[mh] AND CRISPR[tiab]` | 2020:2026[dp], English | 42 |Search results can change as PubMed adds or edits records. Always save the search date.
Final Check
- Are all field tags valid?
- Are
AND,OR, andNOTuppercase? - Are mixed Boolean terms inside parentheses?
- Does a new topic use both MeSH and plain words?
- Is the date range clear?
- Did PubMed map each term as planned?
- Did you avoid filters that may hide good papers?
- Can another person repeat the search from the log?
- Does the code check HTTP errors before reading data?
- Does it handle empty results, timeouts, and rate limits?
- Are API keys read from the environment?
- Are duplicate PMIDs removed?
- Were retractions and corrections checked?