Middleware
Extend functionality with caching, logging, retry, and validation.
Middleware lets you extend and modify the behavior of embedding models and vector databases. There are two kinds, and each built-in middleware belongs to exactly one of them:
| Kind | Wrap with | Built-in middleware |
|---|---|---|
| Embedding model middleware | wrapEmbeddingModel({ model, middleware }) | retryMiddleware, rateLimitMiddleware, piiRedactionMiddleware, dpEmbeddingMiddleware |
| VectorDB middleware | wrapVectorDB({ db, middleware }) | cachingMiddleware, loggingMiddleware, validationMiddleware, encryptionMiddleware |
See it in action
Try the PII Redactor block for a working demo of these APIs.
Embedding Model Middleware
Wrap an embedding model with a single middleware object. To stack several, combine them with composeEmbeddingMiddleware() (see Middleware Composition):
import { wrapEmbeddingModel, retryMiddleware } from '@localmode/core';
import { transformers } from '@localmode/transformers';
const baseModel = transformers.embedding('Xenova/bge-small-en-v1.5');
const model = wrapEmbeddingModel({
model: baseModel,
middleware: retryMiddleware({ maxRetries: 3 }),
});For differentially private embeddings, see dpEmbeddingMiddleware in Differential Privacy.
Vector DB Middleware
Wrap a vector database with one middleware object or an array of them. The wrapped database has the same VectorDB interface as the original:
import { createVectorDB, wrapVectorDB } from '@localmode/core';
const baseDB = await createVectorDB({ name: 'db', dimensions: 384 });
const db = wrapVectorDB({
db: baseDB,
middleware: {
beforeAdd: async (doc) => {
console.log('Adding', doc.id);
return doc;
},
beforeSearch: async (query, options) => {
console.log('Searching with k =', options.k);
return { query, options };
},
afterSearch: async (results) => {
console.log('Found', results.length, 'results');
return results;
},
beforeDelete: async (id) => {
console.log('Deleting', id);
return true; // return false to cancel the delete
},
},
});Vector DB Middleware Interface
Every hook is optional:
interface VectorDBMiddleware {
beforeAdd?: (document: Document) => Document | Promise<Document>;
afterAdd?: (document: Document) => void | Promise<void>;
wrapGet?: (options: {
doGet: () => Promise<Document | null>;
id: string;
}) => Promise<Document | null>;
afterGet?: (document: Document | undefined) => Document | undefined | Promise<Document | undefined>;
beforeDelete?: (id: string) => boolean | Promise<boolean>; // false cancels
afterDelete?: (id: string) => void | Promise<void>;
beforeSearch?: (
query: Float32Array,
options: SearchOptions
) => { query: Float32Array; options: SearchOptions } | Promise<{ query: Float32Array; options: SearchOptions }>;
wrapSearch?: (options: {
doSearch: () => Promise<SearchResult[]>;
query: Float32Array;
options: SearchOptions;
}) => Promise<SearchResult[]>;
afterSearch?: (results: SearchResult[]) => SearchResult[] | Promise<SearchResult[]>;
afterUpdate?: (id: string) => void | Promise<void>;
afterDeleteWhere?: (deletedCount: number) => void | Promise<void>;
afterImport?: () => void | Promise<void>;
beforeClear?: () => boolean | Promise<boolean>; // false cancels
afterClear?: () => void | Promise<void>;
onError?: (error: Error, operation: string) => boolean | void | Promise<boolean | void>;
}wrapSearch and wrapGet wrap the operation itself: call doSearch() / doGet() to run it (through any inner middleware), or return a value without calling it to answer from elsewhere, such as a cache. For a search the order is beforeSearch → wrapSearch → afterSearch; for a get it is wrapGet → afterGet. afterUpdate, afterDeleteWhere, and afterImport fire after update(), deleteWhere(), and import() complete, which is where a middleware that keeps derived state (like the caching middleware) invalidates it.
Error Handling
onError receives every error thrown by a wrapped operation (add, addMany, get, update, delete, deleteMany, deleteWhere, search, clear, import) together with the operation name. Return true to suppress the error: the operation then resolves with a neutral value instead of rejecting. Return false or nothing to let the original error propagate.
| Operation | Resolves with when suppressed |
|---|---|
get | null |
search | [] |
deleteWhere | 0 |
add, addMany, update, delete, deleteMany, clear, import | undefined |
const db = wrapVectorDB({
db: baseDB,
middleware: {
onError: (error, operation) => {
console.error(`VectorDB ${operation} failed:`, error);
return operation === 'search'; // a failed search yields [], every other failure still throws
},
},
});With several middleware, the first one whose onError returns true suppresses the error. The built-in loggingMiddleware logs errors and never suppresses them.
Middleware Composition
Combine multiple middleware into a single middleware using composition functions:
import { composeEmbeddingMiddleware, composeVectorDBMiddleware } from '@localmode/core';Composing Embedding Middleware
import {
composeEmbeddingMiddleware,
piiRedactionMiddleware,
retryMiddleware,
wrapEmbeddingModel,
} from '@localmode/core';
const combined = composeEmbeddingMiddleware([
piiRedactionMiddleware({ emails: true, phones: true }),
retryMiddleware({ maxRetries: 3 }),
]);
const secureModel = wrapEmbeddingModel({ model: baseModel, middleware: combined });Composing VectorDB Middleware
wrapVectorDB accepts an array directly (it composes it for you), or you can compose explicitly with composeVectorDBMiddleware:
import {
composeVectorDBMiddleware,
cachingMiddleware,
loggingMiddleware,
validationMiddleware,
wrapVectorDB,
} from '@localmode/core';
const combined = composeVectorDBMiddleware([
validationMiddleware({ dimensions: 384 }),
loggingMiddleware({ operations: ['search'] }),
cachingMiddleware({ maxSearchResults: 500 }),
]);
const wrappedDB = wrapVectorDB({ db, middleware: combined });Middleware Order
before* and after* hooks run in array order. wrapSearch / wrapGet nest with the first
middleware outermost, so in the example above the logger records every search, including the ones
the cache answers. Put the logger after the cache to log only searches that reach the index.
Factory Aliases
Each built-in middleware also has a create* alias with the same signature and behavior:
| Factory | Same as | Options Type |
|---|---|---|
createCachingMiddleware(options) | cachingMiddleware | CachingMiddlewareOptions |
createLoggingMiddleware(options) | loggingMiddleware | LoggingMiddlewareOptions |
createValidationMiddleware(options) | validationMiddleware | ValidationMiddlewareOptions |
createRetryMiddleware(options) | retryMiddleware | RetryMiddlewareOptions |
createRateLimitMiddleware(options) | rateLimitMiddleware | RateLimitMiddlewareOptions |
Custom Middleware
Custom Embedding Middleware
import { wrapEmbeddingModel, type EmbeddingModelMiddleware } from '@localmode/core';
function timingMiddleware(): EmbeddingModelMiddleware {
return {
transformParams: async ({ values }) => {
// Transform input values
return { values: values.map((v) => v.trim()) };
},
wrapEmbed: async ({ doEmbed, values, model }) => {
const start = Date.now();
// Call the actual embedding function
const result = await doEmbed();
console.log(`Embedded ${values.length} values with ${model.modelId} in ${Date.now() - start}ms`);
return result;
},
};
}
const model = wrapEmbeddingModel({ model: baseModel, middleware: timingMiddleware() });Custom VectorDB Middleware
Implement the VectorDBMiddleware interface to create custom middleware:
import type { VectorDBMiddleware } from '@localmode/core';
function analyticsMiddleware(): VectorDBMiddleware {
let searchCount = 0;
let addCount = 0;
return {
wrapSearch: async ({ doSearch, options }) => {
searchCount++;
const start = performance.now();
const results = await doSearch();
console.log(
`Search #${searchCount} (k=${options.k ?? 10}): ${results.length} results in ${performance.now() - start}ms`
);
return results;
},
afterAdd: async () => {
addCount++;
console.log(`Total documents added: ${addCount}`);
},
};
}Best Practices
Middleware Tips
- Use the right wrapper - Embedding middleware goes to
wrapEmbeddingModel, VectorDB middleware towrapVectorDB - Order matters - Validation first; decide whether logging sits outside or inside the cache
- Keep middleware focused - One concern per middleware
- Write through the wrapper - Middleware only sees operations made on the wrapped instance
- Consider performance - Each middleware adds overhead
Next Steps
Blocks
| App | Description | Links |
|---|---|---|
| Privacy (PII Redactor) | Differential privacy middleware for embedding models | Live block · Source |