The Rogue Scholar science blog archive has now archived all its currently 55,730 blog posts as PDF files. These PDF files use the archival PDF/A-3A format and include the full-text HTML blog post with metadata embedded as Schema.org attachment, and their PDF URLs are included in the Crossref metadata.

Blog posts are published as web pages in HTML format and often authored in Markdown format. So why PDF? I have been struggling with PDF for a long time, as it is a horrible authoring format, made famous by Peter Murray-Rust:

PDF is a hamburger, and we're trying to turn it back into a cow.

At the same time PDF is an excellent archiving format. The reference manager Papers pioneered the idea of managing reference metadata together with PDF attachments, and other reference managers did something similar. This reference manager – PDF integration was further enhanced when more scholarly content became available as Open Access and reference manager Zotero added a Find Available PDF integration.

At the same time PDF/A became an official ISO standard for long-term preservation of PDF documents, making it an archiving option for digital repositories such as Rogue Scholar and complementing web-archiving standards such as WARC developed by the Internet Archive, which Rogue Scholar uses since October 2023.

PDF/A comes in several flavors and Rogue Scholar uses PDF/A-3A that allows attachments of files in any format, in this case the full-text HTML with embedded schema.org metadata. These metadata are embedded similar how this is done on web pages crawled by Google and others, but is added from the blog post metadata on PDF generation. This makes extraction of metadata and full-text HTML from PDF files much easier and faster and doesn't require machine learning tools such as GROBID. The latest versions of the commonmeta-py library can open these PDF files and extract the full-text and metadata, and for example convert to other formats.

The link to the PDF file is included in the metadata of every Rogue Scholar DOI, enabling downstream workflows to automatically fetch the PDF (all blog posts in Rogue Scholar are open access with a CC-BY license). One of the important use cases is of course reference managers, and the Zotero browser integration will now automatically not only fetch metadata from Rogue Scholar, but also the full-text PDF:

Another interesting use case of these archival PDF files is blog platform migration, e.g. from WordPress to Eleventy, or from Blogger to Quarto. While import scripts from other popular platforms exist, this generic archiving format should make it easier to migrate to any blogging platform. More work is of course needed to write import filters for specific platforms, but this is a big step forward.

Finally, the archival PDFs should make it easier to upload full-text content and metadata to other repository platforms, and the obvious starting point is other repositories also using the InvenioRDM platform.

Please reach out via SlackemailMastodon, or Bluesky if you have any questions or comments

Rogue Scholar is a scholarly infrastructure that is free for all authors and readers. You can support Rogue Scholar with a one-time or recurring donation or by becoming a sponsor.

References

  1. Fenner, M. (2010, December 5). Blogging Beyond the PDF. Front Matter. https://doi.org/10.53731/r294649-6f79289-8cw7p
  2. Fenner, M. (2008, October 3). Papers: Interview with Alexander Griekspoor. Front Matter. https://doi.org/10.53731/fy8ppqw-wctn03p
  3. Fenner, M. (2009, March 15). Reference Manager Overview. Front Matter. https://doi.org/10.53731/r294649-6f79289-8cw42
  4. Fenner, M. (2023, October 30). Starting November, all Rogue Scholar blog posts will be archived by the Internet Archive. Front Matter. https://doi.org/10.53731/hhtx0-wb293