Local-first and privacy-focused web utilities are having a massive comeback. With browser engines becoming faster and WebAssembly/Web Workers maturing, there is rarely a reason to push sensitive user documents to an external backend for simple conversions. While building MD-Convert (a zero-upload document to Markdown converter), I explored how to parse real-world documents into clean Markdown entirely on the client side. Here is a breakdown of the core architecture and libraries that make purely in-browser document processing possible. 1. Converting Web Articles with Readability + Turndown Converting messy web markup into clean Markdown involves two distinct steps: Content Extraction: Stripping ads, navbars, sidebars, and trackers. HTML-to-Markdown Transformation: Translating semantic DOM nodes into markdown tokens. Mozilla’s @mozilla/readability paired with turndown is an incredible combination for this: javascript import { Readability } from '@mozilla/readability'; import TurndownService from 'turndown';
function htmlToCleanMarkdown(rawHtmlDocument, sourceUrl) { // 1. Extract pure article content const reader = new Readability(rawHtmlDocument); const article = reader.parse();
if (!article || !article.content) { throw new Error('Unable to extract main content'); }
// 2. Initialize Turndown const turndownService = new TurndownService({ headingStyle: 'atx', codeBlockStyle: 'fenced' });
// Ensure image URLs remain absolute
turndownService.addRule('absoluteImages', {
filter: 'img',
replacement: (content, node) => {
const src = node.getAttribute('src');
const alt = node.getAttribute('alt') || '';
if (!src) return '';
try {
const absoluteUrl = new URL(src, sourceUrl).href;
return \n\n;
} catch {
return \n\n;
}
}
});
return turndownService.turndown(article.content); }
- Offloading Heavy PDF Parsing to Web Workers Parsing large PDFs using pdf.js on the main thread is a recipe for UI freezes and dropped frames. The solution is running the extraction pipeline inside a dedicated Web Worker. Here is a simplified pattern for extracting selectable text layers asynchronously: // worker.js import * as pdfjsLib from 'pdfjs-dist/legacy/build/pdf';
self.onmessage = async (e) => { const { arrayBuffer } = e.data;
try { const loadingTask = pdfjsLib.getDocument({ data: arrayBuffer }); const pdf = await loadingTask.promise; let fullText = '';
for (let pageNum = 1; pageNum <= pdf.numPages; pageNum++) {
const page = await pdf.getPage(pageNum);
const textContent = await page.getTextContent();
const pageText = textContent.items
.map((item) => item.str)
.join(' ');
fullText += `## Page ${pageNum}\n\n${pageText}\n\n`;
}
self.postMessage({ status: 'success', markdown: fullText });
} catch (error) { self.postMessage({ status: 'error', message: error.message }); } };
The Big Advantages of Client-Side Processing Zero Server Costs: The application can run entirely on static hosting (like Cloudflare Pages or GitHub Pages). Absolute Privacy: User data, API keys, or confidential spreadsheets never cross the network. Instant Latency: Conversions happen in-memory without queue waiting times. Thoughts & Edge Cases? The main challenge with 100% in-browser parsing remains edge cases—like complex multi-column PDF layouts or non-OCR scanned documents. Have you built any client-side/local-first tools recently? What libraries do you prefer for client-side parsing? Let's discuss in the comments! 👇

