Document Loaders
Load and parse documents from various sources for RAG pipelines.
Load documents from various formats (text, JSON, CSV, HTML) for processing in RAG pipelines. All loaders are zero-dependency and run entirely in the browser.
Quick Start
import { loadDocument, loadDocuments } from '@localmode/core';
// Auto-detect format and load
const docs = await loadDocument(myFile);
// Load with explicit type
const csvDocs = await loadDocument(csvFile, { loader: 'csv', textColumn: 'content' });
// Batch load multiple sources
const allDocs = await loadDocuments([file1, file2, file3]);Auto-detection
Without a loader option, loadDocument() asks each built-in loader in turn whether it can load the source and uses the first that says yes, in this order:
- JSON — a
.jsonfile /application/jsonblob, or a string that starts with{or[and parses as a JSON object or array (bracketed prose such as"[Note] ..."falls through). - HTML — a
.html/.htmfile /text/htmlblob, or a string containing<html,<!doctype,<head, or<body(case-insensitive). A bare fragment such as<p>…</p>is not detected as HTML; passloader: 'html'for fragments. - CSV — a
.csvfile /text/csvblob, or a string with a header plus at least one data row where every non-empty row has the same number of (two or more) comma-separated fields. Prose that merely contains commas and line breaks is not treated as CSV. - Text — everything else (plain-text files, empty-type blobs, and any string).
Pass loader: 'text' | 'json' | 'csv' | 'html' to skip detection. Any loader-specific option can be passed alongside it and is type-checked against the loader options:
const docs = await loadDocument(csvString, { loader: 'csv', textColumn: 'body', idColumn: 'id' });Loader defaults
Every loader accepts its options in the constructor (and the create*Loader() factories pass them through). Constructor options apply to every load() call; options passed to load() override them for that call:
import { CSVLoader } from '@localmode/core';
const loader = new CSVLoader({ textColumn: 'content', idColumn: 'id' });
const docs = await loader.load(csvFile); // uses textColumn: 'content'
const titles = await loader.load(csvFile, { textColumn: 'title' }); // per-call option winsBuilt-in Loaders
TextLoader
Load plain text content:
import { TextLoader, createTextLoader } from '@localmode/core';
const loader = new TextLoader();
const docs = await loader.load('Hello world');
// Or with default options (same as new TextLoader({ ... }))
const loader2 = createTextLoader({ trim: true, separator: '\n\n' });| Option | Type | Default | Description |
|---|---|---|---|
separator | string | - | Split text by separator into multiple documents |
trim | boolean | true | Trim whitespace from content |
All loaders also accept the shared options:
| Option | Type | Default | Description |
|---|---|---|---|
encoding | string | 'utf-8' | Text encoding |
maxSize | number | - | Maximum file size in bytes |
generateId | (source, index) => string | - | Custom document ID generator |
abortSignal | AbortSignal | - | Cancellation signal |
JSONLoader
Extract text from JSON structures:
import { JSONLoader, createJSONLoader } from '@localmode/core';
const loader = new JSONLoader();
const docs = await loader.load(jsonBlob);
// Extract specific fields
const loader2 = createJSONLoader({
textFields: ['title', 'body'],
fieldSeparator: '\n\n',
recordsPath: 'data.articles',
});| Option | Type | Default | Description |
|---|---|---|---|
textFields | string[] | - | Fields to extract text from |
extractAllStrings | boolean | false | Extract from all string fields |
fieldSeparator | string | '\n' | Separator when combining fields |
recordsPath | string | - | Path to array of records (e.g., 'data.items') |
CSVLoader
Load CSV/TSV data:
import { CSVLoader, createCSVLoader } from '@localmode/core';
const loader = createCSVLoader({
textColumn: 'content',
idColumn: 'id',
hasHeader: true,
});
const docs = await loader.load(csvFile);| Option | Type | Default | Description |
|---|---|---|---|
textColumn | string | number | - | Column for text content |
textColumns | (string | number)[] | - | Multiple columns to combine |
columnSeparator | string | ' ' | Separator for combined columns |
idColumn | string | number | - | Column for document IDs |
columnDelimiter | string | ',' | Column delimiter |
rowDelimiter | string | '\n' | Row delimiter |
hasHeader | boolean | true | First row is header |
skipEmpty | boolean | true | Skip empty rows |
HTMLLoader
Extract text from HTML content:
import { HTMLLoader, createHTMLLoader } from '@localmode/core';
const loader = createHTMLLoader({
selector: 'article',
extractMetadata: true,
ignoreTags: ['script', 'style', 'nav'],
});
const docs = await loader.load(htmlString);| Option | Type | Default | Description |
|---|---|---|---|
selector | string | - | CSS selector to extract from |
selectors | string[] | - | Multiple selectors |
extractMetadata | boolean | true | Extract metadata from <head> |
preserveFormatting | boolean | false | Keep line breaks between blocks instead of collapsing all whitespace to single spaces |
ignoreTags | string[] | script, style, noscript, … | Tags whose content is dropped |
In the browser (DOMParser path) block elements such as <p>, <div>, <li>, <h1>–<h6>, <tr>, and <td> start a new line, so <p>a</p><p>b</p> never merges into ab. With preserveFormatting: true each block stays on its own line; otherwise whitespace is collapsed and blocks are separated by a single space.
DocumentLoader Interface
Implement custom loaders for any format:
import type { DocumentLoader, LoaderSource, LoadedDocument, LoaderOptions } from '@localmode/core';
class MyCustomLoader implements DocumentLoader {
readonly supports = ['.custom', 'application/x-custom'];
canLoad(source: LoaderSource): boolean {
if (source instanceof File) {
return source.name.endsWith('.custom');
}
return false;
}
async load(source: LoaderSource, options?: LoaderOptions): Promise<LoadedDocument[]> {
const text = /* parse your format */;
return [{
id: crypto.randomUUID(),
text,
metadata: { source: 'custom-file', mimeType: 'application/x-custom' },
}];
}
}LoadedDocument
interface LoadedDocument {
id: string;
text: string;
metadata: {
source: string;
mimeType?: string;
title?: string;
pageCount?: number;
sizeBytes?: number;
[key: string]: unknown;
};
}Loader Registry
Create a custom registry with your own loaders for automatic format detection. The registry uses the first loader whose canLoad() accepts the source, so list the most specific loaders first; TextLoader accepts every string and belongs last:
import { createLoaderRegistry, TextLoader, JSONLoader, CSVLoader } from '@localmode/core';
const registry = createLoaderRegistry([
new MyCustomLoader(), // Your custom loader
new JSONLoader(),
new CSVLoader({ textColumn: 'content' }),
new TextLoader(),
]);
// Auto-detect format
const docs = await registry.load(someFile);
// Batch load
const allDocs = await registry.loadMany([file1, file2, file3]);
// Check which loader handles a source
const loader = registry.getLoader(someFile);Accepted Source Types
A LoaderSource is one of:
| Source Type | Description |
|---|---|
string | Raw text content |
Blob | Binary blob |
ArrayBuffer | Raw binary data (decoded with encoding) |
File | Browser File object (from input or drag-and-drop) |
{ type: 'url', url: string } | URL to fetch. No built-in loader's canLoad() claims a URL, so loadDocument() loads it as text unless you pass loader |
{ type: 'custom', data: unknown } | Data for a custom loader. The built-in loaders throw on it |