LocalMode
PDF.js

Overview

PDF text extraction for document processing pipelines.

@localmode/pdfjs

PDF text extraction using PDF.js for local document processing. Extract text, metadata, and structure from PDFs entirely in the browser.

See it in action

Try the Knowledge Base block — it extracts PDF text on-device, then chunks, embeds, and searches it for grounded Q&A.

Features

  • 📄 Full PDF Support — Extract text from any PDF document
  • 🔒 Password Protected — Handle encrypted PDFs
  • 📑 Page-Level Control — Process specific pages or split by page
  • 📊 Metadata Extraction — Get title, author, dates, etc.

Installation

bash pnpm install @localmode/pdfjs @localmode/core
bash npm install @localmode/pdfjs @localmode/core
bash yarn add @localmode/pdfjs @localmode/core
bash bun add @localmode/pdfjs @localmode/core

Quick Start

import { extractPDFText } from '@localmode/pdfjs';

// From file input
const file = document.getElementById('fileInput').files[0];

const { text, pageCount, metadata } = await extractPDFText(file);

console.log(`Extracted ${pageCount} pages`);
console.log('Title:', metadata?.title);
console.log('Text:', text);

API Reference

extractPDFText()

Extract text from a PDF. Accepts File, Blob, ArrayBuffer, Uint8Array, or a URL string as the source:

import { extractPDFText } from '@localmode/pdfjs';

const result = await extractPDFText(source, {
  maxPages: 10, // Limit pages to extract
  includePageNumbers: true, // Add [Page N] headers
  pageSeparator: '\n---\n', // Separator between pages
  password: 'secret', // For encrypted PDFs
});

console.log(result.text); // Full extracted text
console.log(result.pageCount); // Total number of pages
console.log(result.pages); // Array of { pageNumber, text } objects
console.log(result.metadata); // PDF metadata
console.log(result.metadataError); // Set only if the metadata could not be read

Options

Prop

Type

Return Value

Prop

Type

PDFLoader

A DocumentLoader implementation for PDFs. Call its load() method directly: core's loadDocument() only auto-detects the built-in text, JSON, HTML, and CSV loaders, so it does not route PDFs to PDFLoader.

import { PDFLoader } from '@localmode/pdfjs';

const loader = new PDFLoader({
  splitByPage: false, // Single doc or one per page
  maxPages: undefined, // All pages
  includePageNumbers: true,
  password: undefined,
});

// load() returns LoadedDocument[] ({ id, text, metadata })
const documents = await loader.load(pdfBlob);

for (const doc of documents) {
  console.log(doc.text);
  console.log(doc.metadata);
}

Each document's metadata carries:

FieldDescription
sourceFile name, URL, or a generated blob-N / document-N id
mimeType'application/pdf'
pageCountTotal pages in the PDF
titleThe PDF's /Title, if set
createdAtThe PDF's /CreationDate, if set and valid
metadataErrorPresent only when the PDF's information dictionary could not be read (the same reason as PDFExtractResult.metadataError); absent otherwise
page, totalPagesSet on each document when splitByPage is true

load() accepts a File, Blob, ArrayBuffer, a URL string, or { type: 'url', url } (URLs are fetched), and honors abortSignal and generateId from its options argument.

Split by Page

Create separate documents for each page:

import { PDFLoader } from '@localmode/pdfjs';

const loader = new PDFLoader({ splitByPage: true });
const documents = await loader.load(pdfBlob);

console.log(`Loaded ${documents.length} pages`);

for (const doc of documents) {
  // metadata.page and metadata.totalPages are set when splitByPage is true
  console.log(`Page ${doc.metadata.page}: ${doc.text.substring(0, 100)}...`);
}

Utility Functions

import { getPDFPageCount, isPDF } from '@localmode/pdfjs';

// Get page count without full extraction
const pageCount = await getPDFPageCount(pdfBlob);
console.log(`PDF has ${pageCount} pages`);

// Check if file is a PDF
if (await isPDF(file)) {
  // Process as PDF
} else {
  // Handle other file types
}

RAG Pipeline Integration

Build a PDF-powered RAG system:

import { PDFLoader } from '@localmode/pdfjs';
import { createVectorDB, ingest, semanticSearch, streamText } from '@localmode/core';
import { transformers } from '@localmode/transformers';
import { webllm } from '@localmode/webllm';

// Setup
const embeddingModel = transformers.embedding('Xenova/bge-small-en-v1.5');
const llm = webllm.languageModel('Llama-3.2-1B-Instruct-q4f16_1-MLC');
const db = await createVectorDB({ name: 'pdf-docs', dimensions: 384 });

// Load and process PDF
async function ingestPDF(file: File) {
  const loader = new PDFLoader({ splitByPage: true });
  const pages = await loader.load(file);

  // ingest() chunks each page, embeds the chunks, and stores them. Every chunk
  // inherits its page's metadata plus sourceDocId, chunkIndex, chunkStart and
  // chunkEnd, so pass whole pages rather than pre-chunked text.
  const { chunksCreated } = await ingest({
    db,
    model: embeddingModel,
    documents: pages.map((page) => ({
      id: page.id,
      text: page.text,
      metadata: { filename: file.name, page: page.metadata.page },
    })),
    chunking: { strategy: 'recursive', size: 512, overlap: 50 },
  });

  return chunksCreated;
}

// Query
async function queryPDF(question: string) {
  const { results } = await semanticSearch({
    db,
    model: embeddingModel,
    query: question,
    k: 3,
  });

  const context = results.map((r) => `[Page ${r.metadata?.page}]\n${r.text}`).join('\n\n');

  const result = await streamText({
    model: llm,
    prompt: `Answer based on the PDF content:

${context}

Question: ${question}

Answer:`,
  });

  return result;
}

File Upload Component

React example:

import { useState } from 'react';
import { extractPDFText } from '@localmode/pdfjs';

function PDFUploader() {
  const [text, setText] = useState('');
  const [loading, setLoading] = useState(false);

  async function handleFile(e: React.ChangeEvent<HTMLInputElement>) {
    const file = e.target.files?.[0];
    if (!file) return;

    setLoading(true);
    try {
      const { text, pageCount } = await extractPDFText(file);
      setText(text);
      console.log(`Extracted ${pageCount} pages`);
    } catch (error) {
      console.error('Failed to extract PDF:', error);
    } finally {
      setLoading(false);
    }
  }

  return (
    <div>
      <input type="file" accept=".pdf" onChange={handleFile} />
      {loading && <p>Extracting text...</p>}
      {text && <pre>{text}</pre>}
    </div>
  );
}

Handling Large PDFs

For large PDFs, process in chunks:

import { extractPDFText, getPDFPageCount } from '@localmode/pdfjs';

async function processLargePDF(file: File) {
  const { text, pages, pageCount } = await extractPDFText(file, {
    maxPages: 50, // Limit pages if needed
  });

  console.log(`Processed ${pages.length} of ${pageCount} pages`);

  // Access individual page text
  for (const page of pages) {
    console.log(`Page ${page.pageNumber}: ${page.text.substring(0, 100)}...`);
  }

  return text;
}

Password-Protected PDFs

import { extractPDFText } from '@localmode/pdfjs';

try {
  const { text } = await extractPDFText(encryptedPDF, {
    password: userProvidedPassword,
  });
  console.log(text);
} catch (error) {
  if (error.message.includes('password')) {
    // Prompt user for password
  }
}

Metadata Extraction

const { metadata, metadataError } = await extractPDFText(file);

if (metadataError) {
  // The information dictionary could not be read; the text is still usable
  console.warn('No metadata:', metadataError);
} else if (metadata) {
  console.log('Title:', metadata.title);
  console.log('Author:', metadata.author);
  console.log('Subject:', metadata.subject);
  console.log('Creator:', metadata.creator);
  console.log('Creation Date:', metadata.creationDate);
  console.log('Modification Date:', metadata.modificationDate);
}

creationDate and modificationDate are parsed from PDF date strings (D:YYYYMMDDHHmmSSOHH'mm'). A Z suffix is UTC and a +HH'mm' / -HH'mm' suffix is applied as an offset from UTC (hours-only offsets and a missing closing apostrophe are accepted). A date with no zone is read as local time, as the PDF specification leaves its zone unspecified. Malformed or impossible dates (a month of 13, February 30) give undefined rather than an invalid or rolled-over Date.

Runtime Support

@localmode/pdfjs depends on pdfjs-dist 6, which sets these floors:

RuntimeMinimum
Chrome125+
Safari18+
Node.js (server-side extraction)22.13+

In the browser the package loads PDF.js's default build and fetches the worker from jsDelivr (pdfjs-dist@<version>/build/pdf.worker.min.mjs) unless you set GlobalWorkerOptions.workerSrc yourself. On Node.js it loads PDF.js's legacy build (pdfjs-dist/legacy/build/pdf.mjs), the build PDF.js supports for Node, and requests useSystemFonts: true so documents that reference non-embedded standard fonts extract the same text as in a browser. A DOM-emulating test environment such as jsdom runs on Node and gets the Node path. The selection is automatic; there is nothing to configure.

Best Practices

PDF Tips

  1. Split by page - Better for RAG; maintains page context
  2. Use page numbers - Include in metadata for citations
  3. Handle errors - Corrupted PDFs, wrong passwords, etc.
  4. Chunk appropriately - 256-512 chars works well for most PDFs
  5. Check file size - Large PDFs may need batched processing

Next Steps

Composed Block

BlockDescriptionLinks
RAG ChatStreaming grounded RAG over your own text/PDF corpus with inline citations, on a core ⇄ LangChain engine toggleLive · Install: npx shadcn add @localmode/ui/blocks/knowledge/rag-chat

On this page