{"id":7028,"date":"2026-08-09T09:33:14","date_gmt":"2026-08-09T09:33:14","guid":{"rendered":"https:\/\/primetoolhub.com\/?p=7028"},"modified":"2026-08-09T09:55:43","modified_gmt":"2026-08-09T09:55:43","slug":"how-pdf-files-work","status":"publish","type":"post","link":"https:\/\/schoolict.net\/tools\/how-pdf-files-work\/","title":{"rendered":"How PDF Files Work Inside &#8211; Text, Pages and Why Redaction Is Hard"},"content":{"rendered":"<div class=\"pth-hero-section\">\n<div class=\"pth-hero-content\">\n<h2>How PDF Files Work Inside<\/h2>\n<p> A plain explanation of what is actually stored inside a PDF, why copied text comes out jumbled, why a black box does not delete anything, and what the file quietly reveals about you<\/p>\n<div id=\"pth-toc-placeholder\"><\/div>\n<\/p>\n<\/div>\n<div class=\"pth-hero-image\">\n    <img data-no-lazy=\"1\"\n         src=\"https:\/\/schoolict.net\/tools\/wp-content\/uploads\/2026\/08\/how-PDF-files-work-800x447.jpeg\"\n         width=\"800\"\n         height=\"447\"\n         alt=\"how PDF files work\"\n         fetchpriority=\"high\"\n         loading=\"eager\"\n         decoding=\"async\"\n         style=\"width:100%; height:auto; display:block;\">\n  <\/div>\n<\/div>\n\n\n<div class=\"wp-block-rank-math-toc-block\" id=\"rank-math-toc\"><h2>Table of Contents<\/h2><nav><ul><li><a href=\"#\ud83d\udfe2-a-pdf-is-not-a-document-it-is-a-set-of-drawing-instructions\">\ud83d\udfe2 A PDF Is Not a Document, It Is a Set of Drawing Instructions<\/a><\/li><li><a href=\"#\ud83d\udfe1-what-is-actually-inside-the-file\">\ud83d\udfe1 What Is Actually Inside the File<\/a><\/li><li><a href=\"#\ud83d\udd34-why-extracted-text-comes-out-jumbled\">\ud83d\udd34 Why Extracted Text Comes Out Jumbled<\/a><ul><li><a href=\"#the-three-failures-you-will-meet\">The three failures you will meet<\/a><\/li><\/ul><\/li><li><a href=\"#\ud83d\udd34-why-a-black-box-deletes-nothing\">\ud83d\udd34 Why a Black Box Deletes Nothing<\/a><ul><li><a href=\"#what-genuinely-removes-text\">What genuinely removes text<\/a><\/li><\/ul><\/li><li><a href=\"#\ud83d\udfe1-the-part-of-the-file-you-did-not-write\">\ud83d\udfe1 The Part of the File You Did Not Write<\/a><\/li><li><a href=\"#\ud83d\udfe2-why-doing-this-work-in-the-browser-changes-the-picture\">\ud83d\udfe2 Why Doing This Work in the Browser Changes the Picture<\/a><\/li><li><a href=\"#\ud83d\udfe2-practical-rules-worth-remembering\">\ud83d\udfe2 Practical Rules Worth Remembering<\/a><\/li><li><a href=\"#\ud83d\udfe1-keep-reading\">\ud83d\udfe1 Keep Reading<\/a><\/li><\/ul><\/nav><\/div>\n\n\n\n<p class=\"wp-block-paragraph\">Last Updated: August 2026<\/p>\n\n\n\n<h2 id=\"\ud83d\udfe2-a-pdf-is-not-a-document-it-is-a-set-of-drawing-instructions\" class=\"wp-block-heading\">\ud83d\udfe2 A PDF Is Not a Document, It Is a Set of Drawing Instructions<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Nearly every problem people hit with PDFs comes from one wrong assumption: that the file contains a document the way a Word file does, with paragraphs, headings and a flow of text. It does not. Adobe designed the format in the early 1990s to solve a printing problem &#8211; a page had to look identical on every machine and every printer. The answer was to stop describing content and start describing appearance.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What a page actually holds is closer to a script for a plotter. Move to this coordinate. Select this font at this size. Draw these glyphs. Move again. Draw a line from here to there. The result looks like a page of prose to you, but the file has no idea it is prose. If you want the background, the&nbsp;<a href=\"https:\/\/en.wikipedia.org\/wiki\/PDF\" rel=\"noreferrer noopener\" target=\"_blank\">format&#8217;s history and specification<\/a>&nbsp;is a reasonable starting point.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Once that clicks, the odd behaviour stops being mysterious.<\/p>\n\n\n\n<h2 id=\"\ud83d\udfe1-what-is-actually-inside-the-file\" class=\"wp-block-heading\">\ud83d\udfe1 What Is Actually Inside the File<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">A PDF is built from numbered objects that reference each other. Open one in a text editor and you will see fragments like this near the top:<\/p>\n\n\n\n<pre class=\"wp-block-preformatted\">%PDF-1.7\n1 0 obj\n  &lt;&lt; \/Type \/Catalog \/Pages 2 0 R &gt;&gt;\nendobj\n2 0 obj\n  &lt;&lt; \/Type \/Pages \/Kids [3 0 R] \/Count 1 &gt;&gt;\nendobj\n3 0 obj\n  &lt;&lt; \/Type \/Page \/Parent 2 0 R \/Contents 4 0 R &gt;&gt;\nendobj<\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Object 1 is the catalogue, the entry point. It points at object 2, the page tree. That points at object 3, a single page. The page points at object 4, its content stream &#8211; the compressed list of drawing commands that produces what you see.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Four kinds of object matter for everyday work:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Object<\/th><th>What it holds<\/th><th>Why you notice it<\/th><\/tr><\/thead><tbody><tr><td>Page tree<\/td><td>The order of pages<\/td><td>Reordering pages rewrites this, not the pages themselves<\/td><\/tr><tr><td>Content stream<\/td><td>Drawing commands for one page<\/td><td>This is where text and shapes actually live<\/td><\/tr><tr><td>Resources<\/td><td>Fonts, embedded images, colour spaces<\/td><td>Why extracting images is separate from extracting text<\/td><\/tr><tr><td>Info dictionary<\/td><td>Author, title, producer, dates<\/td><td>The metadata that travels with the file<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Because pages are entries in a tree, reordering or deleting them is a cheap operation. The page content is untouched &#8211; only the list that names them changes. That is why page organisation is fast even on a large document, while redaction is slow.<\/p>\n\n\n\n<h2 id=\"\ud83d\udd34-why-extracted-text-comes-out-jumbled\" class=\"wp-block-heading\">\ud83d\udd34 Why Extracted Text Comes Out Jumbled<\/h2>\n\n\n<figure class=\"pth-article-figure pth-img-left\" style=\"float:left; width:700px; max-width:100%; margin:4px 28px 16px 0; clear:left;\"><img decoding=\"async\" src=\"https:\/\/schoolict.net\/tools\/wp-content\/uploads\/2026\/08\/pdf-text-extraction-columns-800x447.jpeg\" alt=\" pdf-text-extraction-columns\" width=\"700\" height=\"394\" loading=\"lazy\" data-no-lazy=\"1\" class=\"pth-article-img\" style=\"width:100%;height:auto;display:block;border-radius:10px;border:1px solid #e2e8f0;\"><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">A content stream places text with commands roughly like&nbsp;<code>BT \/F1 12 Tf 72 720 Td (Hello) Tj ET<\/code>. Read that as: begin text, use font F1 at 12 points, move to coordinate 72,720, show the string &#8220;Hello&#8221;, end text.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Notice what is missing. There is no instruction saying &#8220;this is a paragraph&#8221; or &#8220;this line ends here&#8221;. A line of prose might be a single string, or it might be twenty separate placements because the producing program adjusted spacing between words. Two columns are simply text placed at different x coordinates, with no marker saying they are columns.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So when software extracts text, it has to reconstruct the reading order by looking at coordinates &#8211; grouping fragments that share a similar vertical position into a line, guessing where a space belongs from the gap between glyphs. It works well on a plain report and falls apart on a two-column academic paper, a table, or a form.<\/p>\n\n\n\n<h4 id=\"the-three-failures-you-will-meet\" class=\"wp-block-heading\">The three failures you will meet<\/h4>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\ud83d\udd35&nbsp;<strong>Columns interleaved.<\/strong>&nbsp;The extractor reads across the page instead of down each column, mixing two sentences together.<\/li>\n\n\n\n<li>\ud83d\udfe0&nbsp;<strong>Missing or extra spaces.<\/strong>&nbsp;Word gaps are inferred from distance, so tight kerning produces&nbsp;<code>runtogetherwords<\/code>&nbsp;and loose spacing splits a word.<\/li>\n\n\n\n<li>\ud83d\udfe3&nbsp;<strong>Nothing at all.<\/strong>&nbsp;If the page is a scan, there are no text commands to read &#8211; only one large image. Getting words out needs optical character recognition, a completely different process.<\/li>\n<\/ul>\n\n\n\n<h2 id=\"\ud83d\udd34-why-a-black-box-deletes-nothing\" class=\"wp-block-heading\">\ud83d\udd34 Why a Black Box Deletes Nothing<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This is the most costly misunderstanding in the whole format, and it has caused real leaks in court filings and government releases.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A content stream is a sequence. Commands run in order, and later commands paint over earlier ones. When you draw a black rectangle across a name, you append one more instruction to the end of that sequence. The command that drew the name is still sitting there, earlier in the stream, completely intact.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Anything that reads the stream rather than the picture &#8211; a text extractor, a search index, or simply selecting with your mouse and pressing copy &#8211; walks straight past the rectangle and finds the original characters.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The same trap applies to a white rectangle, a solid image pasted on top, and to page cropping. Cropping changes the visible boundary of the page. The content outside that boundary is still in the file and reappears the moment someone widens the crop.<\/p>\n\n\n\n<h3 id=\"what-genuinely-removes-text\" class=\"wp-block-heading\">What genuinely removes text<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">There are two honest approaches, and both have a cost.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The first is surgical: parse the content stream, find the text-showing commands that fall inside the marked area, and rewrite the stream without them. This keeps the rest of the page selectable, but it is delicate work &#8211; a partially overlapping word, text drawn as a path rather than glyphs, or an unusual font encoding can all leave fragments behind.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The second is blunt and reliable: render the page to an image at print resolution, paint the black bars onto that image, and replace the page with the flattened picture. Nothing can be recovered because no text commands survive. The trade-off is that the whole page stops being selectable or searchable. When the information is genuinely sensitive, that trade is usually the right one, and it is the approach used in the&nbsp;<a href=\"https:\/\/schoolict.net\/tools\/free-offline-pdf-reader-and-editor\/\">offline PDF editor<\/a>&nbsp;on this site.<\/p>\n\n\n\n<h2 id=\"\ud83d\udfe1-the-part-of-the-file-you-did-not-write\" class=\"wp-block-heading\">\ud83d\udfe1 The Part of the File You Did Not Write<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Beyond the visible page, a PDF carries an information dictionary and often an XMP metadata block. Between them they typically record the author name taken from your operating system account, the application that created the file, the exact creation and modification timestamps, and sometimes the original file path.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">None of this appears on the page. All of it travels with the file when you email it. For a public tender, an anonymous submission, or a document leaving a company, that hidden detail is worth clearing deliberately rather than hoping nobody looks.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">There is a second, subtler issue. PDF supports incremental saving, where an edit is appended to the end of the file rather than rewritten from scratch. Done carelessly, the earlier version of an edited page can remain inside the file, recoverable by anyone who reads the raw objects. Saving a fresh copy rather than repeatedly patching the same file avoids that.<\/p>\n\n\n\n<h2 id=\"\ud83d\udfe2-why-doing-this-work-in-the-browser-changes-the-picture\" class=\"wp-block-heading\">\ud83d\udfe2 Why Doing This Work in the Browser Changes the Picture<\/h2>\n\n\n\n<div style=\"float: left; width: 48%; min-width: 300px; margin-right: 20px; margin-bottom: 15px;\">\n    <div class=\"pth-inline-card\" data-url=\"\/free-offline-pdf-reader-and-editor\/\"><\/div>\n<\/div>\n\n\n\n<p class=\"wp-block-paragraph\">Every one of the operations above &#8211; parsing objects, walking the page tree, rendering a page to pixels, writing a new file &#8211; can now run inside a browser tab. Modern browsers expose the file to a page as raw bytes through the&nbsp;<a href=\"https:\/\/developer.mozilla.org\/en-US\/docs\/Web\/API\/File_API\" rel=\"noreferrer noopener\" target=\"_blank\">File API<\/a>, and the same rendering engine that draws this article can draw a PDF page onto a canvas.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The practical consequence is a privacy one. A document containing a salary figure, a medical result or a signature never has to be handed to a third party just so it can be annotated. It is opened, changed and saved on the machine it was already sitting on. For documents that carry personal detail, that is a meaningfully different risk profile from a free upload site, and it is the same reasoning behind the wider set of tools described in our guide to&nbsp;<a href=\"https:\/\/schoolict.net\/tools\/secure-offline-web-development-utilities-guide\/\">secure offline web development utilities<\/a>.<\/p>\n\n\n\n<h2 id=\"\ud83d\udfe2-practical-rules-worth-remembering\" class=\"wp-block-heading\">\ud83d\udfe2 Practical Rules Worth Remembering<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li>\ud83d\udd35 If the text must be gone, flatten the page to an image. Covering it is decoration, not removal.<\/li>\n\n\n\n<li>\ud83d\udfe0 Test your own redaction the way an opponent would: save the file, reopen it, select the area and press copy.<\/li>\n\n\n\n<li>\ud83d\udfe3 Clear the metadata before a document leaves your organisation.<\/li>\n\n\n\n<li>\ud83d\udd35 Expect messy extraction from multi-column layouts and nothing at all from scans.<\/li>\n\n\n\n<li>\ud83d\udfe0 Save a new copy rather than repeatedly patching the same file, so old revisions do not travel along.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Try it on a real document<\/strong><br>Annotate, reorder, sign and permanently redact a PDF without uploading it anywhere.<br><a href=\"https:\/\/schoolict.net\/tools\/free-offline-pdf-reader-and-editor\/\">Open PDF Studio Pro<\/a><\/p>\n\n\n\n<h2 id=\"\ud83d\udfe1-keep-reading\" class=\"wp-block-heading\">\ud83d\udfe1 Keep Reading<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If you need to combine several PDFs or cut one apart, that job belongs to the\u00a0<a href=\"https:\/\/schoolict.net\/tools\/pdf-merge-and-split-studio\/\">PDF Merge &amp; Split Studio<\/a>. If you would rather produce a clean PDF from scratch, writing in Markdown and exporting is often faster than fighting a word processor &#8211; the method is covered in\u00a0<a href=\"https:\/\/schoolict.net\/tools\/offline-markdown-to-pdf-converter-article\/\">how Markdown converts to PDF<\/a>. For a wider view of what runs without uploading anything, see the\u00a0<a href=\"https:\/\/schoolict.net\/tools\/top-10-client-side-offline-web-developer-tools\/\">round-up of client-side offline tools<\/a>\u00a0or browse the full\u00a0<a href=\"https:\/\/schoolict.net\/tools\/free-web-tools-directory-prime-tool-hub\/\">free web tools directory<\/a>.<\/p>\n\n\n\n<style>\n.pth-faq-section{margin:50px auto 40px;font-family:inherit;max-width:1480px;padding:0 20px;box-sizing:border-box}\n.pth-faq-header{font-size:1.8rem;font-weight:800;color:#0f172a;margin-bottom:25px;border-bottom:2px solid #e2e8f0;padding-bottom:10px;display:flex;align-items:center;gap:10px}\n.pth-faq-grid{display:grid;grid-template-columns:1fr;gap:20px}\n@media(min-width:768px){.pth-faq-grid{grid-template-columns:repeat(2,1fr)}}\n@media(min-width:1024px){.pth-faq-grid{grid-template-columns:repeat(3,1fr)}}\n.pth-faq-card{background:#f8fafc;padding:24px;border-radius:12px;border:1px solid #e2e8f0;transition:transform .2s ease;break-inside:avoid}\n.pth-faq-card:hover{transform:translateY(-3px);box-shadow:0 4px 12px rgba(0,0,0,.05)}\n.pth-faq-q{color:#0f172a;font-size:1rem;font-weight:700;margin:0 0 12px;line-height:1.4}\n.pth-faq-a{margin:0;font-size:.95rem;color:#1e293b;line-height:1.6;font-weight:500}\n<\/style>\n \n<div class=\"pth-faq-wrap\">\n  <p class=\"pth-faq-title\">Frequently Asked Questions<\/p>\n  <div class=\"pth-faq-grid\">\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">What is stored inside a PDF file?<\/p>\n      <p class=\"pth-faq-a\">Numbered objects that reference each other: a catalogue, a page tree, a content stream of drawing commands for each page, embedded fonts and images, and a dictionary of file information.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">Why does copied PDF text look scrambled?<\/p>\n      <p class=\"pth-faq-a\">The file stores glyph positions, not lines or paragraphs. Software has to rebuild the reading order from coordinates, which struggles with columns, tables and unusual spacing.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">Can text under a black box really be recovered?<\/p>\n      <p class=\"pth-faq-a\">Yes. The rectangle is a later drawing command painted over earlier ones. The original text command is still in the content stream and can be read by anything that parses the file.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">Does cropping a page delete the hidden part?<\/p>\n      <p class=\"pth-faq-a\">No. Cropping changes the visible page boundary only. The content outside it remains in the file and returns as soon as the crop is widened.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">What is the safest way to redact?<\/p>\n      <p class=\"pth-faq-a\">Render the page to an image, paint the bars onto that image, and replace the page with it. No text commands survive, so nothing can be recovered. The page stops being selectable, which is the trade-off.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">Why can I not extract text from a scan?<\/p>\n      <p class=\"pth-faq-a\">A scanned page holds one large image rather than text commands. Turning those pixels into characters requires optical character recognition, which is a separate process from extraction.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">What personal information does a PDF carry?<\/p>\n      <p class=\"pth-faq-a\">Commonly the author name from your user account, the software that produced the file, creation and modification timestamps, and sometimes the original file path. None of it shows on the page.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">Why is reordering pages fast but redaction slow?<\/p>\n      <p class=\"pth-faq-a\">Reordering only rewrites the page tree, a small list. Redaction has to render a page to pixels at print resolution and rebuild it, which is far more work.<\/p>\n    <\/div>\n \n    <div class=\"pth-faq-card\">\n      <p class=\"pth-faq-q\">Can a PDF still contain an older version of an edit?<\/p>\n      <p class=\"pth-faq-a\">It can. PDF allows incremental saving, where changes are appended rather than rewritten. Saving a fresh copy instead of patching the same file repeatedly avoids carrying old revisions along.<\/p>\n    <\/div>\n \n  <\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>How PDF Files Work Inside A plain explanation of what is actually stored inside a PDF, why copied text comes out jumbled, why a black box does not delete anything, and what the file quietly reveals about you Last Updated: August 2026 \ud83d\udfe2 A PDF Is Not a Document, It Is a Set of Drawing &#8230; <a title=\"How PDF Files Work Inside &#8211; Text, Pages and Why Redaction Is Hard\" class=\"read-more\" href=\"https:\/\/schoolict.net\/tools\/how-pdf-files-work\/\" aria-label=\"Read more about How PDF Files Work Inside &#8211; Text, Pages and Why Redaction Is Hard\">Read more<\/a><\/p>\n","protected":false},"author":1,"featured_media":7030,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[25],"tags":[],"class_list":["post-7028","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-document-tool"],"_links":{"self":[{"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/posts\/7028","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/comments?post=7028"}],"version-history":[{"count":4,"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/posts\/7028\/revisions"}],"predecessor-version":[{"id":7124,"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/posts\/7028\/revisions\/7124"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/media\/7030"}],"wp:attachment":[{"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/media?parent=7028"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/categories?post=7028"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/schoolict.net\/tools\/wp-json\/wp\/v2\/tags?post=7028"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}