Your PDF and its HTML version carry the same material — which one should rank, and how do you tell Google without breaking the other? This deck is a release checklist that shows how to prevent PDF duplicate content in search: define the duplicate set, choose the preferred version, and send that decision through HTTP signals crawlers can read.
Inside: the canonical Link header for PDF files, X-Robots-Tag rules that block indexing by accident, Yandex text-layer and file-size limits, and a final sign-off list for duplicate URL consolidation before release. Each slide is one item to verify, with the response fields and status codes to check, plus edge cases the standard guides skip — crawl-blocked directives, image-only scans, and conflicting canonical targets.
When the cleanup ships, confirm which URL search engines picked with the Bulk Index Checker by SpeedyIndex: https://en.speedyindex.com/google-index-checker/