LiteRT-LM Web API

The Web API of LiteRT-LM for JavaScript and TypeScript in the browser.

Supported Models

The LiteRT-LM JS API supports the following web-compatible models:

For vector embedding generation in the browser (such as EmbeddingGemma 2 using EmbeddingEngine), see the Embedding Models guide.

Introduction

Here is a sample REPL chat app built with the JavaScript API:

<div id="out" style="white-space: pre-wrap; font-family: monospace;"></div>
<input id="in" onkeydown="if(event.key === 'Enter') repl(this)">

<script type="module">
  import { Engine } from 'https://cdn.jsdelivr.net/npm/@litert-lm/core/+esm';
  const engine = await Engine.create({
    // Load the Gemma 4 E2B model
    model: 'https://huggingface.co/litert-community/gemma-4-E2B-it-litert-lm/resolve/main/gemma-4-E2B-it-gpu.litertlm'
    // Or use the E4B model by swapping in this line
    // model: 'https://huggingface.co/litert-community/gemma-4-E4B-it-litert-lm/resolve/main/gemma-4-E4B-it-gpu.litertlm'
  });
  const chat = await engine.createConversation();

  window.repl = async (el) => {
    const text = el.value;
    el.value = ''; // Clear immediately
    out.append(`\n>>> ${text}\nAI: `);

    for await (const chunk of chat.sendMessageStreaming(text)) {
      out.append(chunk.content[0].text);
    }
  };
</script>

Getting Started

LiteRT-LM is available as an npm package. You can install the latest version from npm or directly import it from a CDN:

# From npm
npm i --save @litert-lm/core

# From a CDN (in your JavaScript file)
import * as litertlm from 'https://cdn.jsdelivr.net/npm/@litert-lm/core/+esm';

Initialize the Engine

The Engine is the entry point to the API. It handles model loading, session creation, and resource management. Remember to delete the engine to release resources when the model is no longer needed.

Note: Initializing the engine can take several seconds to load the model.

import {Engine, EngineSettings} from '@litert-lm/core';

const engineSettings = {
  model: 'url/path/to/model.litertlm', // or a ReadableStream, or a Blob

  // You can configure context length and other settings here
  mainExecutorSettings: {
    maxNumTokens: 8192,
  },
} satisfies EngineSettings;

const engine = await Engine.create(engineSettings);

// ... Use the engine to create a conversation ...

// Delete the engine when done.
await engine.delete();

Flexible Model Sources & Streaming

The model field supports multiple input types:

  • URL string: Fetches and streams weights directly over HTTP or HTTPS.
  • Blob or File: Loads models from local storage, IndexedDB, or the File System Access API.
  • ReadableStream<Uint8Array>: Supports custom progressive streaming pipelines.
// Example: Loading from a Blob or local File handle
const fileInput = document.querySelector('input[type="file"]') as HTMLInputElement;
const file = fileInput.files![0];

const engine = await Engine.create({
  model: file,
});

Custom WASM Binary Location

If you are self-hosting WASM files or deploying in offline environments, load the WASM runtime before calling Engine.create():

import {loadLiteRtLm} from '@litert-lm/core';

// Load from a custom directory containing the WASM binaries
await loadLiteRtLm('/path/to/wasm/directory/');

Create a Conversation

Once the engine is initialized, create a Conversation instance. You can provide a ConversationConfig to customize system messages and sampling parameters.

import {ConversationConfig, SamplerType} from '@litert-lm/core';

const config: ConversationConfig = {
  // Initial background messages and instructions
  preface: {
    messages: [
      {role: 'system', content: 'You are a helpful and concise coding assistant.'},
    ],
  },
  // Generation and sampling controls
  sessionConfig: {
    maxOutputTokens: 2048,
    samplerParams: {
      type: SamplerType.TOP_P,
      temperature: 0.7,
      p: 0.9,
      k: 40,
    },
  },
  // Optional optimizations
  enableConstrainedDecoding: true,
  prefillPrefaceOnInit: true,
  filterChannelContentFromKvCache: true,
};

const conversation = await engine.createConversation(config);

Send Messages

You can send messages with or without streaming.

Non-Streaming Example

// Simple string input
let response = await conversation.sendMessage("What is the capital of France?");
console.log(response.content[0].text);

// Or with full message structure
response = await conversation.sendMessage({role: 'user', content: '...'});

Streaming Example

// sendMessageStreaming returns a ReadableStream of response chunks
const stream = conversation.sendMessageStreaming('Tell me a long story.');

for await (const chunk of stream) {
  // Chunks are Records containing pieces of the response
  for (const item of chunk.content) {
    if (item.type === 'text') {
      console.log(item.text);
    }
  }
}

Cancel Generation

You can cancel an ongoing generation explicitly by calling cancel() on the Conversation instance:

// Cancel any ongoing generation
conversation.cancel();

If you are streaming the response, exiting the for await...of loop early (such as with break) will also automatically cancel the ongoing generation:

for await (const chunk of stream) {
  if (shouldStop()) {
    break; // Cancels the stream and underlying generation
  }
}

Manage Conversation State

Conversation provides methods for branching, inspecting history, and monitoring tokens:

// 1. Fork a conversation into an independent branch without re-prefilling history
const branchedChat = await conversation.clone();

// 2. Inspect the current token count of the conversation context
const tokenCount = await conversation.getTokenCount();
console.log(`Tokens in context: ${tokenCount}`);

// 3. Retrieve conversation history
const history = await conversation.getHistory();
console.log('History:', history);

// 4. Retrieve inference performance benchmarks
const benchmarkInfo = await conversation.getBenchmarkInfo();
console.log(`Prefill rate: ${benchmarkInfo.lastPrefillTokensPerSecond} tokens/sec`);
console.log(`Decode rate: ${benchmarkInfo.lastDecodeTokensPerSecond} tokens/sec`);

Thinking

When running reasoning-capable models, the LiteRT-LM Web SDK allows generating internal reasoning tokens ("thinking") before returning the final response.

You can toggle thinking mode by setting enable_thinking in extra_context within the conversation preface:

import { Engine } from '@litert-lm/core';

// 1. Initialize the Engine
const engine = await Engine.create({
  model: 'url/path/to/thinking_model.litertlm'
});

// 2. Configure thinking using extra_context in the conversation preface
const conversation = await engine.createConversation({
  preface: {
    extra_context: {
      enable_thinking: true, // Set to false to disable thinking
    }
  },
  sessionConfig: {
    maxOutputTokens: 2048,
  }
});

// 3. Send a message
const response = await conversation.sendMessage("Solve this logic puzzle step by step.");

// 4. Access reasoning tokens and final answer
if (response.channels?.['thought']) {
  console.log("Thought:", response.channels['thought']);
}
console.log("Answer:", response.content?.[0]?.text ?? response.content);

// Clean up
await engine.delete();