LocalMode
LangChain

LangChain

Drop-in local inference for existing LangChain.js applications.

LangChain Integration

@localmode/langchain provides adapter classes so existing LangChain.js applications can swap cloud providers for 100% local inference — by changing just 3 imports.

See it in action

Try the Knowledge Base block — its LangChain engine toggle runs LocalModeEmbeddings, LocalModeVectorStore, and ChatLocalMode end-to-end in the browser, behind the same UI as the core pipeline.

Adapters

LangChain Base ClassLocalMode AdapterWrapsProvider
EmbeddingsLocalModeEmbeddingsEmbeddingModel@localmode/transformers
BaseChatModelChatLocalModeLanguageModel@localmode/webllm
VectorStoreLocalModeVectorStoreVectorDB@localmode/core
BaseDocumentCompressorLocalModeRerankerRerankerModel@localmode/transformers

Each adapter is a thin, stateless wrapper — it converts between LangChain and LocalMode data formats and delegates all work to the underlying model or database.

Installation

pnpm install @localmode/langchain @localmode/core @localmode/transformers
# Add webllm if using ChatLocalMode:
pnpm install @localmode/webllm

Supported @langchain/core versions

@localmode/langchain accepts @langchain/core 0.3.80 or later on the 0.3 line, and 1.1.0 or later on the 1.x line. Versions below 0.3.80 have a serialization-injection flaw that can leak secrets, and 1.0.0 through 1.0.6 shipped a dist/ without the top-level entry points that moduleResolution: "node" projects need. The adapters extend Embeddings, VectorStore, BaseChatModel, and BaseDocumentCompressor, which are identical across both supported lines, so the adapter examples on these pages run unchanged on either. (Classes from the separate langchain package can move between its majors; see the ContextualCompressionRetriever note in the Migration Guide.) If your app already depends on @langchain/core, make sure it resolves inside that range so a single copy is installed.

Quick Start — Full RAG Chain

import { LocalModeEmbeddings, ChatLocalMode, LocalModeVectorStore } from '@localmode/langchain';
import { transformers } from '@localmode/transformers';
import { webllm } from '@localmode/webllm';
import { createVectorDB } from '@localmode/core';

// 1. Create local models
const embeddingModel = transformers.embedding('Xenova/bge-small-en-v1.5');
const llmModel = webllm.languageModel('Qwen3-1.7B-q4f16_1-MLC');

// 2. Wrap in LangChain adapters
const embeddings = new LocalModeEmbeddings({ model: embeddingModel });
const llm = new ChatLocalMode({ model: llmModel });

// 3. Create vector store backed by local IndexedDB
const db = await createVectorDB({ name: 'docs', dimensions: 384 });
const store = new LocalModeVectorStore(embeddings, { db });

// 4. Add documents (embeds automatically)
await store.addDocuments([
  { pageContent: 'LocalMode runs AI in the browser.', metadata: { source: 'docs' } },
  { pageContent: 'Data never leaves the device.', metadata: { source: 'docs' } },
]);

// 5. Search
const results = await store.similaritySearch('privacy', 3);

// 6. Generate with context
const context = results.map((r) => r.pageContent).join('\n');
const answer = await llm.invoke(`Based on: ${context}\n\nQuestion: How does LocalMode handle privacy?`);

Everything runs locally. No API keys, no servers, no data leaves the device.

Knowledge Base Engine

createLangChainKnowledgeBaseEngine() returns a kind: 'langchain' engine that implements the same frozen KnowledgeBaseEngine contract exported from @localmode/core (chunk → embed → store, vector search, grounded ask) — but runs it through the real LocalModeEmbeddings, LocalModeVectorStore, and ChatLocalMode adapters over an in-memory createVectorDB. It exists so an app can toggle the whole pipeline between the core implementation and the LangChain adapters without changing a line of calling code.

import { createLangChainKnowledgeBaseEngine, ChatLocalMode } from '@localmode/langchain';
import { transformers } from '@localmode/transformers';

const engine = createLangChainKnowledgeBaseEngine({
  embeddingModel: transformers.embedding('Xenova/bge-small-en-v1.5'),
  getChatModel: () =>
    new ChatLocalMode({
      model: transformers.languageModel('onnx-community/granite-4.0-350m-ONNX-web'),
      maxTokens: 512,
    }),
});
await engine.ingest(docs, { chunking: 'recursive', chunkSize: 500 });
const hits = await engine.search('privacy and encryption', { topK: 10 });
const { answer, sources } = await engine.ask('How is data encrypted?');

Options

CreateLangChainKnowledgeBaseEngineOptions:

OptionTypeDefaultNotes
embeddingModelEmbeddingModel— (required)Defines the corpus space; wrapped internally in LocalModeEmbeddings.
getChatModel() => Promise<ChatLocalMode> | ChatLocalMode— (required)Lazy — resolved (and awaited) only on the first ask(), so nothing loads on construction. The caller owns model construction, any device pre-probe, and the ChatLocalMode maxTokens / temperature budget.
storage'memory' | 'indexeddb' | StorageAdapter'memory'VectorDB storage backend (in-memory session store by default).
chunkDefaultsLangChainKnowledgeBaseChunkDefaults{ chunkSize: 500, chunkOverlap: 50 }Applied when IngestOptions omits sizes.
askConfigLangChainKnowledgeBaseAskConfig{ topK: 4 }systemPrompt? grounds the answer; topK? is the retrieval depth when AskOptions.topK is omitted.

Because the models are injected, @localmode/langchain gains no provider dependency — apps that never toggle the LangChain engine never pull it.

Equivalence with the core engine

The LangChain engine is result-equivalent to @localmode/core's createKnowledgeBaseEngine: given the same corpus, embedding model, and query, both return the same ranked contract ids and the same cosine scores (higher-is-better, in [0, 1]). To hold that guarantee:

  • Chunking uses a faithful local port of LangChain's RecursiveCharacterTextSplitter (same default separators, keepSeparator, and merge/overlap semantics). 'semantic' mode falls back to recursive — equivalence is defined at the result level, not byte-for-byte chunk parity.
  • Contract ids (${docId}:${chunkIndex}) are reconstructed from chunk metadata, because LocalModeVectorStore generates its own UUIDs for stored vectors.
  • Grounded answers strip <think>…</think> reasoning blocks before returning, matching the core engine.

New exports: createLangChainKnowledgeBaseEngine, CreateLangChainKnowledgeBaseEngineOptions, LangChainKnowledgeBaseAskConfig, LangChainKnowledgeBaseChunkDefaults.

The frozen contract and full KnowledgeBaseEngine API reference live in core — see Core Knowledge Base. For the React session hook, see useKnowledgeBase. The Knowledge Base block toggles the core and LangChain engines live over one shared corpus.

Key Design Decisions

  • User provides model instances — You create the LocalMode model and pass it to the adapter. The adapter doesn't know about provider packages.
  • Float32Array to number[] — LangChain uses number[][] for embeddings. The adapter converts automatically via Array.from().
  • Streaming fallback — ChatLocalMode._stream() uses the model's doStream() if available, otherwise falls back to generating the full response and yielding it as a single chunk.
  • No tool calling — Local models have limited tool-calling ability. ChatLocalMode returns text-only content.

Package Details

PropertyValue
Package@localmode/langchain
Dependencies@langchain/core (>=0.3.80 <1.0.0 || >=1.1.0 <2.0.0)
Peer Dependencies@localmode/core (>=2.3.0)
BundleESM + CJS, tree-shakeable
Side EffectsNone

For individual adapter docs, see: Embeddings, Chat Model, Vector Store, Migration Guide.

Composed Block

BlockDescriptionLinks
RAG ChatStreaming grounded RAG over your own text/PDF corpus with inline citations, on a core ⇄ LangChain engine toggleLive · Install: npx shadcn add @localmode/ui/blocks/knowledge/rag-chat

On this page