The Web API of LiteRT-LM for JavaScript and TypeScript in the browser.
Supported Models
The LiteRT-LM JS API supports the following web-compatible models:
- Gemma 4 E2B
(
gemma-4-E2B-it-gpu.litertlm) from litert-community/gemma-4-E2B-it-litert-lm - Gemma 4 E4B
(
gemma-4-E4B-it-gpu.litertlm) from litert-community/gemma-4-E4B-it-litert-lm - Gemma 4 12B
(
gemma-4-12B-it-gpu.litertlm) from litert-community/gemma-4-12B-it-litert-lm - Gemma 4 26B A4B
(
gemma-4-26B-A4B-it-gpu.litertlm) from litert-community/gemma-4-26B-A4B-it-litert-lm - Gemma 4 31B
(
gemma-4-31B-it-gpu.litertlm) from litert-community/gemma-4-31B-it-litert-lm
For vector embedding generation in the browser (such as EmbeddingGemma 2 using
EmbeddingEngine), see the Embedding Models guide.
Introduction
Here is a sample REPL chat app built with the JavaScript API:
<div id="out" style="white-space: pre-wrap; font-family: monospace;"></div>
<input id="in" onkeydown="if(event.key === 'Enter') repl(this)">
<script type="module">
import { Engine } from 'https://cdn.jsdelivr.net/npm/@litert-lm/core/+esm';
const engine = await Engine.create({
// Load the Gemma 4 E2B model
model: 'https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm/resolve/main/gemma-4-E2B-it-gpu.litertlm'
// Or use the E4B model by swapping in this line
// model: 'https://huggingface.co/litert-community/gemma-4-E4B-it-litert-lm/resolve/main/gemma-4-E4B-it-gpu.litertlm'
});
const chat = await engine.createConversation();
window.repl = async (el) => {
const text = el.value;
el.value = ''; // Clear immediately
out.append(`\n>>> ${text}\nAI: `);
for await (const chunk of chat.sendMessageStreaming(text)) {
out.append(chunk.content[0].text);
}
};
</script>
Getting Started
LiteRT-LM is available as an npm package. You can install the latest version from npm or directly import it from a CDN:
# From npm
npm i --save @litert-lm/core
# From a CDN (in your JavaScript file)
import * as litertlm from 'https://cdn.jsdelivr.net/npm/@litert-lm/core/+esm';
Initialize the Engine
The Engine is the entry point to the API. It handles model loading, session
creation, and resource management. Remember to delete the engine to release
resources when the model is no longer needed.
Note: Initializing the engine can take several seconds to load the model.
import {Engine, EngineSettings} from '@litert-lm/core';
const engineSettings = {
model: 'url/path/to/model.litertlm', // or a ReadableStream, or a Blob
// You can configure context length and other settings here
mainExecutorSettings: {
maxNumTokens: 8192,
},
} satisfies EngineSettings;
const engine = await Engine.create(engineSettings);
// ... Use the engine to create a conversation ...
// Delete the engine when done.
await engine.delete();
Flexible Model Sources & Streaming
The model field supports multiple input types:
- URL string: Fetches and streams weights directly over HTTP or HTTPS.
BloborFile: Loads models from local storage, IndexedDB, or the File System Access API.ReadableStream<Uint8Array>: Supports custom progressive streaming pipelines.
// Example: Loading from a Blob or local File handle
const fileInput = document.querySelector('input[type="file"]') as HTMLInputElement;
const file = fileInput.files![0];
const engine = await Engine.create({
model: file,
});
Custom WASM Binary Location
If you are self-hosting WASM files or deploying in offline environments, load
the WASM runtime before calling Engine.create():
import {loadLiteRtLm} from '@litert-lm/core';
// Load from a custom directory containing the WASM binaries
await loadLiteRtLm('/path/to/wasm/directory/');
Create a Conversation
Once the engine is initialized, create a Conversation instance. You can
provide a ConversationConfig to customize system messages and sampling
parameters.
import {ConversationConfig, SamplerType} from '@litert-lm/core';
const config: ConversationConfig = {
// Initial background messages and instructions
preface: {
messages: [
{role: 'system', content: 'You are a helpful and concise coding assistant.'},
],
},
// Generation and sampling controls
sessionConfig: {
maxOutputTokens: 2048,
samplerParams: {
type: SamplerType.TOP_P,
temperature: 0.7,
p: 0.9,
k: 40,
},
},
// Optional optimizations
enableConstrainedDecoding: true,
prefillPrefaceOnInit: true,
filterChannelContentFromKvCache: true,
};
const conversation = await engine.createConversation(config);
Send Messages
You can send messages with or without streaming.
Non-Streaming Example
// Simple string input
let response = await conversation.sendMessage("What is the capital of France?");
console.log(response.content[0].text);
// Or with full message structure
response = await conversation.sendMessage({role: 'user', content: '...'});
Streaming Example
// sendMessageStreaming returns a ReadableStream of response chunks
const stream = conversation.sendMessageStreaming('Tell me a long story.');
for await (const chunk of stream) {
// Chunks are Records containing pieces of the response
for (const item of chunk.content) {
if (item.type === 'text') {
console.log(item.text);
}
}
}
Cancel Generation
You can cancel an ongoing generation explicitly by calling cancel() on the
Conversation instance:
// Cancel any ongoing generation
conversation.cancel();
If you are streaming the response, exiting the for await...of loop early (such
as with break) will also automatically cancel the ongoing generation:
for await (const chunk of stream) {
if (shouldStop()) {
break; // Cancels the stream and underlying generation
}
}
Manage Conversation State
Conversation provides methods for branching, inspecting history, and
monitoring tokens:
// 1. Fork a conversation into an independent branch without re-prefilling history
const branchedChat = await conversation.clone();
// 2. Inspect the current token count of the conversation context
const tokenCount = await conversation.getTokenCount();
console.log(`Tokens in context: ${tokenCount}`);
// 3. Retrieve conversation history
const history = await conversation.getHistory();
console.log('History:', history);
// 4. Retrieve inference performance benchmarks
const benchmarkInfo = await conversation.getBenchmarkInfo();
console.log(`Prefill rate: ${benchmarkInfo.lastPrefillTokensPerSecond} tokens/sec`);
console.log(`Decode rate: ${benchmarkInfo.lastDecodeTokensPerSecond} tokens/sec`);
Thinking
When running reasoning-capable models, the LiteRT-LM Web SDK allows generating internal reasoning tokens ("thinking") before returning the final response.
You can toggle thinking mode by setting enable_thinking in extra_context
within the conversation preface:
import { Engine } from '@litert-lm/core';
// 1. Initialize the Engine
const engine = await Engine.create({
model: 'url/path/to/thinking_model.litertlm'
});
// 2. Configure thinking using extra_context in the conversation preface
const conversation = await engine.createConversation({
preface: {
extra_context: {
enable_thinking: true, // Set to false to disable thinking
}
},
sessionConfig: {
maxOutputTokens: 2048,
}
});
// 3. Send a message
const response = await conversation.sendMessage("Solve this logic puzzle step by step.");
// 4. Access reasoning tokens and final answer
if (response.channels?.['thought']) {
console.log("Thought:", response.channels['thought']);
}
console.log("Answer:", response.content?.[0]?.text ?? response.content);
// Clean up
await engine.delete();