Skip to content

On-Device AI (Gemma)

Pitcher exposes AI inference to apps through one JS-API surface — the app calls a method and doesn't care where the answer comes from. Impact owns the model on each platform:

  • iOS runs a Gemma model fully on-device through llama.cpp with Metal acceleration. Inference is offline, private (nothing leaves the device), and free per query.
  • Web has no local model — the same surface routes to the online (Bedrock) backend instead.

Your app talks to it through usePitcherApi(), exactly like any other bridge call.

js
import { usePitcherApi } from '@pitcher/js-api'
const api = usePitcherApi()

Which method should I use?

MethodUse it whenPlatforms
piaSearchAnswer()You want an answer to a rep's question over the instance's content. Start here — it handles on-device vs online routing for you.iOS + Web
aiGetCapabilities()You need to know whether/what AI is available before showing UI.iOS + Web
aiComplete()You need raw, low-level text completion on the on-device model.iOS only (web rejects)

TIP

Most apps only need piaSearchAnswer(). Reach for aiComplete() only when you're driving the on-device model directly and are handling the iOS-only / web-fallback split yourself.

Field casing follows your SDK casing option

Responses (and raw ai.* event bodies) come back in snake_case by default (casing: 'snake'). If you initialize the API with casing: 'camel', they are camelCased instead — result.tokens_generated becomes result.tokensGenerated. Request payload keys are already snake_case (max_tokens) — don't rename them. The examples below use the default snake_case.

piaSearchAnswer() — the high-level call

Answers a natural-language question over the instance's content and returns a unified result, regardless of which runtime served it.

js
const result = await api.piaSearchAnswer({
  query: 'What is our latest pricing for oncology products?',
})

console.log(result.answer)          // the answer text
console.log(result.source)          // 'on_device' | 'online' — which runtime answered
console.log(result.cited_file_ids)  // files the answer drew on (online path)
console.log(result.citations)       // which PAGES of those files (online path, flag-gated)
console.log(result.confidence)      // 'high' | 'medium' | 'low' (online path)

Page citations

When the instance has page-level citations enabled and the answer came from document text, the result additionally carries citations — one entry per cited file, with the pages the answer was grounded in:

js
// result.citations -> [{ file_id: '01ABC…', page_indices: [1, 4] }]
for (const citation of result.citations ?? []) {
  for (const pageIndex of citation.page_indices) {
    // page_indices are 0-BASED — pass them straight to an open-at-page call,
    // and add 1 only for a user-facing label ("Page 2").
    renderPageChip(citation.file_id, pageIndex, `Page ${pageIndex + 1}`)
  }
}

citations is optional at runtime: the key is absent — never null, never [] — when the feature is off for the instance, when the file has no page information, or when the app is running an older CDN build of the SDK. The answer route guarantees that cited_file_ids is a superset of citations[].file_id — it builds that list as the union of the two — so a file with no page information still shows as a source. The SDK does not re-derive the relation, so treat it as a backend guarantee rather than something the client enforces. Always guard with result.citations ?? [] and fall back to a file-level citation.

Payload (PiaSearchAnswerPayload):

  • query (required) — the rep's question.
  • instance_id — defaults to the env's active instance.
  • context_text — client-assembled content context (file names/tags/summaries). Used only by the on-device path; the online route builds its own context server-side.
  • file_ids — narrows the online route's candidate pool.
  • max_tokens — token budget for the on-device completion.

Routing is driven by the pia_search_config LaunchDarkly flag (see Configuration):

  • on_device and iOS and a model is available → on-device Gemma over your context_text. Any failure (bridge timeout, model unavailable/downloading, completion error) transparently falls back to online.
  • online (or any on-device fallback) → the next-core Bedrock route, which assembles candidate context server-side.
  • off → the call rejects (pia_search_disabled).

Because fallback is automatic, a single piaSearchAnswer() call is safe to make on any platform — you only need to handle the answer and a possible rejection.

aiGetCapabilities() — is AI available?

js
const caps = await api.aiGetCapabilities()
if (!caps.available) {
  // Hide AI UI. caps.unavailability_reason explains why (iOS).
}
  • On web it resolves statically to the online runtime: { available: true, offline: false, runtime: 'bedrock-online' }.
  • On iOS it reflects the native on-device runtime (AIRuntimeCapabilities):
    • available — inference is possible right now.
    • offline — true for the on-device path.
    • llm_model — the loaded model identifier.
    • model_available_for_download, model_downloading, download_progress_percent — download state, so you can render "Tap to enable AI" / "Downloading 42%".
    • unavailability_reason — human-readable "why not" when available is false.

aiComplete() — raw on-device completion (iOS only)

Low-level, text-in / text-out completion on the on-device model. iOS only — on web the host rejects; use piaSearchAnswer() for online inference.

js
const completion = await api.aiComplete({
  prompt: 'Summarize the following call notes in 3 bullets:\n' + notes,
  max_tokens: 256, // optional; on-device runtime defaults to 512
})
console.log(completion.text)

The payload is the exact iOS wire contract: { prompt, max_tokens?, stream? }{ text, tokens_generated?, elapsed_seconds?, time_to_first_token_seconds? }.

TIP

The runtime is model-agnostic and applies no chat template. For best results with Gemma, wrap your prompt yourself:

js
const prompt = `<start_of_turn>user\n${userText}<end_of_turn>\n<start_of_turn>model\n`

On-device constraints (enforced natively): prompt capped at 32,000 bytes; max_tokens defaults to 512, hard-capped at 1024; a 30s per-request budget; one inference at a time.

Configuration

These are set per instance (backend / LaunchDarkly), not from your app, but they govern behavior:

SettingValues / defaultPurpose
pia_search_config (LD){ mode: 'off' | 'on_device' | 'online' }How piaSearchAnswer() serves inference.
offline_ai_model_tag (iOS)default offline-ai-modelThe Pitcher file tag the Gemma GGUF asset must carry, so the device can find and download it.
offline_ai_cellular_download_policy (iOS)always_ask / wifi_only / allow_cellularWhether to confirm the (multi-GB) model download on cellular.
enable_meeting_notes_ai (iOS)default offGates the on-device meeting-notes capture/summarize pipeline.

The Gemma model file (~2.9 GB) is downloaded on demand the first time it's needed on a device; there's a native confirmation prompt subject to the cellular policy. Until it's present, aiGetCapabilities() reports model_available_for_download / model_downloading.

Advanced: raw bridge access (iOS)

The iOS on-device runtime exposes more than the typed SDK surface above — function calling, cancellation, token streaming, model-download events, and meeting-notes summaries. These are not part of the typed @pitcher/js-api surface; you reach them through the low-level bridge on usePitcherApi(). They are iOS-only and their shapes may change — prefer the typed methods above where they suffice.

Raw requests go through api.API.request(type, body); native events are consumed with api.on(type, cb) / api.off(type, cb).

Function calling (ai.function_call)

An agentic loop over a fixed, native-side Tier-1 CRM tool catalog (backed by the on-device SmartStore). The catalog is not supplied from JS.

js
const res = await api.API.request('ai.function_call', {
  prompt: 'Which of my accounts have opportunities closing this month?',
  // max_tokens_per_turn?: number, stream?: boolean
})
// res: { text, tool_calls: [{ function, arguments, status, result, elapsed_seconds }], iterations, ... }

Token streaming

Set stream: true on ai.complete / ai.function_call to receive incremental tokens. The originating request still resolves with the full text — streaming is additive.

js
const onToken = ({ delta }) => appendToUi(delta) // { request_id, delta, index }
api.on('ai.token', onToken)
api.on('ai.token_stream_completed', () => api.off('ai.token', onToken)) // { request_id } — fires once before the request settles

await api.API.request('ai.complete', { prompt, stream: true })

For ai.function_call, only the model's final natural-language turn streams; intermediate tool-call JSON turns don't.

Cancellation (ai.cancel)

Cancel an in-flight request by its id; it settles with the CANCELLED error code and any partial output is preserved. You obtain the id from the request_id on ai.token events, so cancellation is practical for streamed requests.

js
let activeRequestId
api.on('ai.token', ({ request_id }) => { activeRequestId = request_id })
// ...
if (activeRequestId) api.API.request('ai.cancel', { request_id: activeRequestId })

Model-download events

Fired while the on-device model is fetched on demand:

  • ai.model_download_progress{ bytes_downloaded, bytes_total, percent }
  • ai.model_download_completed{ bytes_total } (safe to retry your request)
  • ai.model_download_failed{ reason }
js
api.on('ai.model_download_progress', ({ percent }) => setProgress(percent))
api.on('ai.model_download_completed', () => retryMyRequest())

Meeting-notes summary

When enable_meeting_notes_ai is on, Impact captures call audio, transcribes it on-device (whisper.cpp), and summarizes it with Gemma — fully offline; audio and transcript never leave the device. Your app just receives the finished summary:

js
api.on('meeting_notes_summary_pending', () => showSpinner())
api.on('meeting_notes_summary_ready', ({ summary }) => {
  // e.g. append into a post-call Notes field (append below; never replace)
  notesField.value += summary
})

This is how the post-call app populates its Notes field.

Error codes

Raw ai.* requests reject with { code, message }. Branch on code:

CodeMeaning
SESSION_ACTIVEAnother inference is already running.
PROMPT_TOO_LARGEPrompt over 32,000 bytes.
TIMEOUTExceeded the 30s budget.
CANCELLEDResult of ai.cancel (usually expected).
INSUFFICIENT_MEMORYDevice under memory pressure.
MODEL_NOT_AVAILABLE / MODEL_NOT_DOWNLOADED / MODEL_DOWNLOADINGModel isn't ready — a download may have been prompted; retry after ai.model_download_completed.
MODEL_DOWNLOAD_DECLINED / MODEL_DOWNLOAD_INSUFFICIENT_STORAGE / MODEL_DOWNLOAD_OFFLINE / MODEL_DOWNLOAD_FAILEDDownload couldn't proceed — surface the matching "tap to retry" / "free up space" / "connect to a network" affordance.
AI_DISABLED / MODEL_NOT_LOADED / MODEL_LOAD_FAILED / DECODE_FAILED / INVALID_REQUESTRuntime-level failures — treat as "AI unavailable".

See also