add Canadian Paediatric Society statement scraper - #9
Conversation
Scrapes the intersection of PubMed's Practice Guideline publication type with the PMC Open Access subset, 2,829 guidelines. That is the only slice where the full text is both retrievable and openly licensed; the other 34k Practice Guideline records expose an abstract only, under publisher copyright. Discovery and extraction use NCBI's E-utilities rather than the rendered pages. esearch paginates by numeric offset, so the page number maps straight onto retstart and an empty page past the end terminates the run. esummary resolves a whole batch of PMIDs to PMCIDs and citation metadata in one request. efetch returns JATS XML carrying body, section structure and license. JATS is close enough to HTML that renaming tags and reusing html_to_markdown is cheaper and less error-prone than a second serializer: table-wrap already contains genuine XHTML tables, and inline markup maps one to one. Licensing is recorded per document rather than claimed for the source, because the Open Access Subset is not uniformly Creative Commons licensed. Censused over all 2,829 records: CC BY 42.8% publisher terms, no CC license 13.3% CC BY-NC 20.5% of which: Elsevier COVID grant 207 CC BY-NC-ND 20.2% PMC OA "unrestricted" 51 CC BY-NC-SA 1.9% no <license> element 101 CC0 1.3% So 44.1% is CC BY or CC0 and carries no restriction on derivatives, 22.4% is non-commercial only, 20.2% asserts NoDerivatives, and 13.3% needs reading case by case. Presence in the subset is not itself a grant to redistribute: 101 records carry only a copyright line such as "(c) Springer-Verlag Tokyo 2007", and Elsevier's pandemic-era deposits grant free access while still reserving all rights. The license name is parsed from the Creative Commons URL rather than the license-type attribute, which the corpus spells 16 different ways. The copyright statement is captured separately because it is a sibling of <license>, not a child, and it holds the reservation of rights. Figures are recorded in metadata rather than linked. Unlike the HTML sources in MedARC-AI#9 and MedARC-AI#10, which absolutize a real <img src>, JATS carries only a bare filename; the served URL inserts a CDN shard and content hash that appear nowhere in the API response, so a constructed link 404s. The filename, label and caption are kept so a later pass can resolve them without re-scraping, and the caption stays in the text. E-utilities calls retry with backoff on 429 and 5xx. One document makes up to two calls back to back and the rate limit is per source address, so NCBI does answer with 429 in practice; without a retry that propagates past the skip handler and kills the whole run. Records PMC holds without a deposited body, 1.1% of the corpus, are logged and skipped rather than aborting the run.
Scrapes the intersection of PubMed's Guideline publication type with the PMC Open Access subset, 3,184 guidelines. That is the only slice where the full text is both retrievable and openly licensed; the other 40k Guideline records expose an abstract only, under publisher copyright. `Guideline[pt]` rather than `Practice Guideline[pt]`: the former is a strict superset, and the 352 records it adds are clinical rather than administrative (the 2025 Korean CPR guidelines and similar), so the narrower tag would drop 11% of the corpus for nothing. Discovery and extraction use NCBI's E-utilities rather than the rendered pages. esearch paginates by numeric offset, so the page number maps straight onto retstart and an empty page past the end terminates the run. esummary resolves a whole batch of PMIDs to PMCIDs and citation metadata in one request. efetch returns JATS XML carrying body, section structure and license. JATS is close enough to HTML that renaming tags and reusing html_to_markdown is cheaper and less error-prone than a second serializer: table-wrap already contains genuine XHTML tables, and inline markup maps one to one. Licensing is recorded per document rather than claimed for the source, because the Open Access Subset is not uniformly Creative Commons licensed. Censused over all 3,184 records: CC BY 43.9% publisher terms, no CC license 12.6% CC BY-NC 21.1% of which: Elsevier COVID grant 211 CC BY-NC-ND 19.2% no <license> element 112 CC BY-NC-SA 2.1% PMC OA "unrestricted" 58 CC0 1.2% other publisher terms 19 By what that permits: 45.1% unrestricted for derivative works, 23.2% non-commercial only, 19.2% asserting NoDerivatives, 12.6% needing a case-by-case reading. Presence in the subset is not itself a grant to redistribute: 112 records carry only a copyright line such as "(c) Springer-Verlag Tokyo 2007", and Elsevier's pandemic-era deposits grant free access while still reserving all rights. The license name is parsed from the Creative Commons URL rather than the license-type attribute, which the corpus spells 18 different ways. The copyright statement is captured separately because it is a sibling of <license>, not a child, and it holds the reservation of rights. Figures and supplementary files are recorded in metadata rather than linked. Unlike the HTML sources in MedARC-AI#9 and MedARC-AI#10, which absolutize a real <img src>, JATS carries only a bare filename; the served URL inserts a CDN shard and content hash that appear nowhere in the API response, so a constructed link 404s. Supplementary blocks are pointers too: across 40 sampled guidelines every one referenced an external .docx or .tif rather than inline content, totalling 0.18% of body text. Recording name, label and caption keeps the evidence tables findable without re-scraping. E-utilities calls retry with backoff on 429 and 5xx. One document makes up to two calls back to back and the rate limit is per source address, so NCBI does answer with 429 in practice; without a retry that propagates past the skip handler and kills the whole run. external_id prefers the PMID from the record itself, so an article reached from a PMC URL gets the same identifier as one reached from the listing. Records PMC holds without a deposited body, 1.0% of the corpus, are logged and skipped rather than aborting the run.
Scrapes the intersection of PubMed's Guideline publication type with the PMC Open Access subset, restricted to English: about 3,000 guidelines. That is the only slice where the full text is both retrievable and openly licensed; the other 40k Guideline records expose an abstract only, under publisher copyright. Two scoping choices, both measured rather than assumed: `Guideline[pt]` rather than `Practice Guideline[pt]`. The former is a strict superset and the 352 records it adds are clinical, not administrative (the 2025 Korean CPR guidelines and similar), so the narrower tag would drop 11% of the corpus for nothing. `English[la]`, which drops 185 records. Every other source in this package is already English-only as a side effect of its entry URL: CPS is a bilingual site scraped through its /en/ routes, WHO publishes in six languages and is scraped through its English listing. PubMed's API returns every language, so the filter has to be explicit to match. The excluded records are largely French CMAJ translations of guidelines already in the corpus. Discovery and extraction use NCBI's E-utilities rather than the rendered pages. esearch paginates by numeric offset, so the page number maps straight onto retstart and an empty page past the end terminates the run. esummary resolves a whole batch of PMIDs to PMCIDs and citation metadata in one request. efetch returns JATS XML carrying body, section structure and license. JATS is close enough to HTML that renaming tags and reusing html_to_markdown is cheaper and less error-prone than a second serializer: table-wrap already contains genuine XHTML tables, and inline markup maps one to one. Licensing is recorded per document rather than claimed for the source, because the Open Access Subset is not uniformly Creative Commons licensed. Censused over all 2,999 English records present on 2026-07-29: CC BY 44.5% publisher terms, no CC license 11.7% CC BY-NC 22.1% of which: Elsevier COVID grant, no CC BY-NC-ND 18.4% <license> element at all (112), CC BY-NC-SA 2.1% PMC OA "unrestricted re-use" CC0 1.3% By what that permits: 45.8% unrestricted for derivative works, 24.1% non-commercial only, 18.4% asserting NoDerivatives, 11.7% needing a case-by-case reading. Presence in the subset is not itself a grant to redistribute: 112 records carry only a copyright line such as "(c) Springer-Verlag Tokyo 2007", and Elsevier's pandemic-era deposits grant free access while still reserving all rights. The license name is parsed from the Creative Commons URL rather than the license-type attribute, which the corpus spells 18 different ways. The copyright statement is captured separately because it is a sibling of <license>, not a child, and it holds the reservation of rights. Figures and supplementary files are recorded in metadata rather than linked. Unlike the HTML sources in MedARC-AI#9 and MedARC-AI#10, which absolutize a real <img src>, JATS carries only a bare filename; the served URL inserts a CDN shard and content hash that appear nowhere in the API response, so a constructed link 404s. Supplementary blocks are pointers too: across 40 sampled guidelines every one referenced an external .docx or .tif rather than inline content, totalling 0.18% of body text. Recording name, label and caption keeps the evidence tables findable without re-scraping. E-utilities calls retry with backoff on 429 and 5xx. One document makes up to two calls back to back and the rate limit is per source address, so NCBI does answer with 429 in practice; without a retry that propagates past the skip handler and kills the whole run. external_id prefers the PMID from the record itself, so an article reached from a PMC URL gets the same identifier as one reached from the listing. Records PMC holds without a deposited body, 0.6% of the corpus, are logged and skipped rather than aborting the run.
| CPS retains copyright in its position statements and practice points. This | ||
| module supports a permission-gated ingestion workflow; obtain written CPS | ||
| permission before running a corpus scrape or redistributing its output. | ||
| """ |
There was a problem hiding this comment.
https://cps.ca/en/policies-politiques/copyright
"While linking to our website is encouraged and welcome, the CPS usually does not allow content to reside on other websites."
There was a problem hiding this comment.
Yes, the content of CPS should not be directly published as part of an open source dataset.
|
|
||
|
|
||
| class _CpsMarkdownConverter(MarkdownConverter): | ||
| """Preserve CPS superscripts that carry clinical meaning or citations.""" |
There was a problem hiding this comment.
Here you have your own custom converter using markdownify rather than using the shared html_to_markdown class. It appears this is because CPS citations would be inappropriately stripped with that shared class?
Perhaps it would be good to implement a different solution that uses the shared converter to reduce maintenance costs in the future?
There was a problem hiding this comment.
I removed the MarkdownConverter subclass, and instead added the non-destructive preserving sup_symbol="<sup>" to the base MarkdownConverter.
I think it's cleaner this way too.
| assert ( | ||
| "Legacy citation[2](https://cps.ca/en/documents/position/febrile-young-infants#ref2)." | ||
| in document.content | ||
| ) |
There was a problem hiding this comment.
These two assertions need to be unwrapped to pass the ruff format checks, which are causing the other tests to be skipped. This can be done with
uv run ruff format datasets && git commit -am "ruff format"
| image.drop_tree() | ||
|
|
||
| for link in root.xpath(".//a[@href]"): | ||
| link.set("href", urljoin(base_url, link.get("href"))) |
There was a problem hiding this comment.
This code rewrites everly link target on the page to an absolute URL. Unfortunately this leads to links to the same page being expanded, and appears to lead to 10% of the scraped text being self-URLs.
FYI the scraped document text appears to be composed of ~56% clinical prose, ~11% self-URLs, and ~32% references.
There was a problem hiding this comment.
Removed absolute URL normalization for same-document references. However, images, other links and assets are still normalized to absolute URLs.
| section_count = max(1, len(root.xpath(".//*[self::h2 or self::h3 or self::h4 or self::h5 or self::h6]"))) | ||
| content = _content_to_markdown(root, link_mode=link_mode) | ||
| if not content: | ||
| raise CpsFetchError("No readable statement content found on the CPS page") |
There was a problem hiding this comment.
This check only flags when markdown is empty.
However, some CPS statements are not published as HTML, but rather PDFs. This leads to some documents only containing an abstract. For example the extracted markdown for https://cps.ca/en/documents/position/covid-19-vaccine-for-children-and-adolescents-technical-report contains only ~700 characters whereas the actual technical report is in a PDF.
Some potential solutions to this:
- record PDF href in metadata and flag when it is the only substantive content
- record the content lenght in the metadata
There was a problem hiding this comment.
I found 127 statements containing PDF links, but only this one seems to have the main material in PDF, whereas the others generally link to supplemental material, cited papers, etc.
As suggested, I still did:
- Record content length in metadata
- Added download_url for available PDFs/other resources
What this does
Adds a Canadian Paediatric Society (CPS) scraper to the datasets scraping pipeline. Same shape as the NICE scraper — discovery and extraction are separate and everything outputs a normalized
ScrapedDocument.The scraper extracts the full HTML position statement or practice point, including tables, references, citation links, and remote figures. Generated assets are not downloaded locally.
How discovery works
Discovery starts from the CPS statements-by-date index at
/en/documents/statements-by-date.CPS uses offset-based pagination in increments of 10 (
/P10,/P20, etc.). Each page is parsed for links under.stmt-title, unrelated links are rejected, and English and language-neutral routes are normalized to canonical HTTPS URLs.The live indexing run discovered and scraped 196 current statements. A 10-second delay is applied between document requests.
How extraction works
Each position statement page is fetched and the primary
.statement-wrappercontent is converted to markdown. There are fallbacks for article and main-content layouts.The extraction preserves:
10<sup>9</sup>/LNavigation, sharing controls, quizzes, related-content blocks, empty links, and decorative PDF icons are removed.
CPS uses a source-specific markdown converter because its legacy and current citation markup needs additional handling. Empty or missing content raises
CpsFetchError.CLI
--source allnow includes CPS alongside NICE.Licensing
CPS statements are not published under an identified open redistribution license. CPS retains copyright and requires permission for corpus scraping or redistribution.
This PR contains scraper code and synthetic test fixtures only. It does not include scraped CPS documents.
QA
Test plan
uv run pytest datasets/test/test_scraping_cps.py— pagination, URL normalization,structured extraction, and error paths
uv run pytest— no regressions (33 passed)uv run ruff check .uv run amfv-scrape --source cps --documents 3 -f markdown -o /tmp/cps-out/