Overview
PDF text extraction for document processing pipelines.
@localmode/pdfjs
PDF text extraction using PDF.js for local document processing. Extract text, metadata, and structure from PDFs entirely in the browser.
See it in action
Try the Knowledge Base block — it extracts PDF text on-device, then chunks, embeds, and searches it for grounded Q&A.
Features
- 📄 Full PDF Support — Extract text from any PDF document
- 🔒 Password Protected — Handle encrypted PDFs
- 📑 Page-Level Control — Process specific pages or split by page
- 📊 Metadata Extraction — Get title, author, dates, etc.
Installation
bash pnpm install @localmode/pdfjs @localmode/core bash npm install @localmode/pdfjs @localmode/core bash yarn add @localmode/pdfjs @localmode/core bash bun add @localmode/pdfjs @localmode/core Quick Start
import { extractPDFText } from '@localmode/pdfjs';
// From file input
const file = document.getElementById('fileInput').files[0];
const { text, pageCount, metadata } = await extractPDFText(file);
console.log(`Extracted ${pageCount} pages`);
console.log('Title:', metadata?.title);
console.log('Text:', text);API Reference
extractPDFText()
Extract text from a PDF. Accepts File, Blob, ArrayBuffer, Uint8Array, or a URL string as the source:
import { extractPDFText } from '@localmode/pdfjs';
const result = await extractPDFText(source, {
maxPages: 10, // Limit pages to extract
includePageNumbers: true, // Add [Page N] headers
pageSeparator: '\n---\n', // Separator between pages
password: 'secret', // For encrypted PDFs
});
console.log(result.text); // Full extracted text
console.log(result.pageCount); // Total number of pages
console.log(result.pages); // Array of { pageNumber, text } objects
console.log(result.metadata); // PDF metadata
console.log(result.metadataError); // Set only if the metadata could not be readOptions
Prop
Type
Return Value
Prop
Type
PDFLoader
A DocumentLoader implementation for PDFs. Call its load() method directly: core's loadDocument() only auto-detects the built-in text, JSON, HTML, and CSV loaders, so it does not route PDFs to PDFLoader.
import { PDFLoader } from '@localmode/pdfjs';
const loader = new PDFLoader({
splitByPage: false, // Single doc or one per page
maxPages: undefined, // All pages
includePageNumbers: true,
password: undefined,
});
// load() returns LoadedDocument[] ({ id, text, metadata })
const documents = await loader.load(pdfBlob);
for (const doc of documents) {
console.log(doc.text);
console.log(doc.metadata);
}Each document's metadata carries:
| Field | Description |
|---|---|
source | File name, URL, or a generated blob-N / document-N id |
mimeType | 'application/pdf' |
pageCount | Total pages in the PDF |
title | The PDF's /Title, if set |
createdAt | The PDF's /CreationDate, if set and valid |
metadataError | Present only when the PDF's information dictionary could not be read (the same reason as PDFExtractResult.metadataError); absent otherwise |
page, totalPages | Set on each document when splitByPage is true |
load() accepts a File, Blob, ArrayBuffer, a URL string, or { type: 'url', url } (URLs are fetched), and honors abortSignal and generateId from its options argument.
Split by Page
Create separate documents for each page:
import { PDFLoader } from '@localmode/pdfjs';
const loader = new PDFLoader({ splitByPage: true });
const documents = await loader.load(pdfBlob);
console.log(`Loaded ${documents.length} pages`);
for (const doc of documents) {
// metadata.page and metadata.totalPages are set when splitByPage is true
console.log(`Page ${doc.metadata.page}: ${doc.text.substring(0, 100)}...`);
}Utility Functions
import { getPDFPageCount, isPDF } from '@localmode/pdfjs';
// Get page count without full extraction
const pageCount = await getPDFPageCount(pdfBlob);
console.log(`PDF has ${pageCount} pages`);
// Check if file is a PDF
if (await isPDF(file)) {
// Process as PDF
} else {
// Handle other file types
}RAG Pipeline Integration
Build a PDF-powered RAG system:
import { PDFLoader } from '@localmode/pdfjs';
import { createVectorDB, ingest, semanticSearch, streamText } from '@localmode/core';
import { transformers } from '@localmode/transformers';
import { webllm } from '@localmode/webllm';
// Setup
const embeddingModel = transformers.embedding('Xenova/bge-small-en-v1.5');
const llm = webllm.languageModel('Llama-3.2-1B-Instruct-q4f16_1-MLC');
const db = await createVectorDB({ name: 'pdf-docs', dimensions: 384 });
// Load and process PDF
async function ingestPDF(file: File) {
const loader = new PDFLoader({ splitByPage: true });
const pages = await loader.load(file);
// ingest() chunks each page, embeds the chunks, and stores them. Every chunk
// inherits its page's metadata plus sourceDocId, chunkIndex, chunkStart and
// chunkEnd, so pass whole pages rather than pre-chunked text.
const { chunksCreated } = await ingest({
db,
model: embeddingModel,
documents: pages.map((page) => ({
id: page.id,
text: page.text,
metadata: { filename: file.name, page: page.metadata.page },
})),
chunking: { strategy: 'recursive', size: 512, overlap: 50 },
});
return chunksCreated;
}
// Query
async function queryPDF(question: string) {
const { results } = await semanticSearch({
db,
model: embeddingModel,
query: question,
k: 3,
});
const context = results.map((r) => `[Page ${r.metadata?.page}]\n${r.text}`).join('\n\n');
const result = await streamText({
model: llm,
prompt: `Answer based on the PDF content:
${context}
Question: ${question}
Answer:`,
});
return result;
}File Upload Component
React example:
import { useState } from 'react';
import { extractPDFText } from '@localmode/pdfjs';
function PDFUploader() {
const [text, setText] = useState('');
const [loading, setLoading] = useState(false);
async function handleFile(e: React.ChangeEvent<HTMLInputElement>) {
const file = e.target.files?.[0];
if (!file) return;
setLoading(true);
try {
const { text, pageCount } = await extractPDFText(file);
setText(text);
console.log(`Extracted ${pageCount} pages`);
} catch (error) {
console.error('Failed to extract PDF:', error);
} finally {
setLoading(false);
}
}
return (
<div>
<input type="file" accept=".pdf" onChange={handleFile} />
{loading && <p>Extracting text...</p>}
{text && <pre>{text}</pre>}
</div>
);
}Handling Large PDFs
For large PDFs, process in chunks:
import { extractPDFText, getPDFPageCount } from '@localmode/pdfjs';
async function processLargePDF(file: File) {
const { text, pages, pageCount } = await extractPDFText(file, {
maxPages: 50, // Limit pages if needed
});
console.log(`Processed ${pages.length} of ${pageCount} pages`);
// Access individual page text
for (const page of pages) {
console.log(`Page ${page.pageNumber}: ${page.text.substring(0, 100)}...`);
}
return text;
}Password-Protected PDFs
import { extractPDFText } from '@localmode/pdfjs';
try {
const { text } = await extractPDFText(encryptedPDF, {
password: userProvidedPassword,
});
console.log(text);
} catch (error) {
if (error.message.includes('password')) {
// Prompt user for password
}
}Metadata Extraction
const { metadata, metadataError } = await extractPDFText(file);
if (metadataError) {
// The information dictionary could not be read; the text is still usable
console.warn('No metadata:', metadataError);
} else if (metadata) {
console.log('Title:', metadata.title);
console.log('Author:', metadata.author);
console.log('Subject:', metadata.subject);
console.log('Creator:', metadata.creator);
console.log('Creation Date:', metadata.creationDate);
console.log('Modification Date:', metadata.modificationDate);
}creationDate and modificationDate are parsed from PDF date strings (D:YYYYMMDDHHmmSSOHH'mm'). A Z suffix is UTC and a +HH'mm' / -HH'mm' suffix is applied as an offset from UTC (hours-only offsets and a missing closing apostrophe are accepted). A date with no zone is read as local time, as the PDF specification leaves its zone unspecified. Malformed or impossible dates (a month of 13, February 30) give undefined rather than an invalid or rolled-over Date.
Runtime Support
@localmode/pdfjs depends on pdfjs-dist 6, which sets these floors:
| Runtime | Minimum |
|---|---|
| Chrome | 125+ |
| Safari | 18+ |
| Node.js (server-side extraction) | 22.13+ |
In the browser the package loads PDF.js's default build and fetches the worker from jsDelivr (pdfjs-dist@<version>/build/pdf.worker.min.mjs) unless you set GlobalWorkerOptions.workerSrc yourself. On Node.js it loads PDF.js's legacy build (pdfjs-dist/legacy/build/pdf.mjs), the build PDF.js supports for Node, and requests useSystemFonts: true so documents that reference non-embedded standard fonts extract the same text as in a browser. A DOM-emulating test environment such as jsdom runs on Node and gets the Node path. The selection is automatic; there is nothing to configure.
Best Practices
PDF Tips
- Split by page - Better for RAG; maintains page context
- Use page numbers - Include in metadata for citations
- Handle errors - Corrupted PDFs, wrong passwords, etc.
- Chunk appropriately - 256-512 chars works well for most PDFs
- Check file size - Large PDFs may need batched processing
Next Steps
Composed Block
| Block | Description | Links |
|---|---|---|
| RAG Chat | Streaming grounded RAG over your own text/PDF corpus with inline citations, on a core ⇄ LangChain engine toggle | Live · Install: npx shadcn add @localmode/ui/blocks/knowledge/rag-chat |
LocalMode Bench
The open cross-runtime browser-AI benchmark - LLM and embedding inference measured across WebLLM, wllama, Transformers.js, LiteRT, and Chrome Built-in AI on real consumer hardware, with a public leaderboard and a CC0 open dataset.
Overview
Enhanced IndexedDB storage with Dexie.js — schema versioning and transactions.