SG Systems

Consulting & forward-deployed engineering for Southeast Asian applied AI.

We embed senior engineers with client teams and ship working systems in weeks rather than producing strategy decks. Fixed scope, outcome-based pricing, no hourly billing.

Serving banking, insurance, fintech, healthcare, retail, government and mid-market commerce across SEA — where a deployment has to be PDPA-compliant, locally tuned, and native to the channels people actually use.

sg-systems.co

Why we publish inference work

Regional deployment runs into two constraints that a general-purpose API does not solve.

Data residency. PDPA and its neighbours make "send the customer's data to a US endpoint" a compliance conversation before it is an engineering one. Inference that never leaves the device sidesteps that conversation entirely.

Hardware reality. SEA deployments run on the devices people have, not on datacentre accelerators. A model that needs 15 GB is not a product; one that fits in 3.7 GB is.

So we work on making capable models run locally. What we publish here are the artifacts and measurements from that work.

Projects

FORGE — mixed-precision post-training quantization. The result depends less on the average bit width than on which tensor families are excluded from it: crushing a Mamba-2 state projection to ternary yields a model that is fluent and confidently wrong.

HELIX — a Metal prefill kernel for the SSM scan in Mamba-2 hybrids, integrated into llama.cpp through a 42-line patch. Chunk-parallel decomposition; ~3x on the isolated op for shapes upstream's fast path declines.

repository
forge quantization
helix Metal kernel + llama.cpp patch
helix-chat-ui SwiftUI client
forge-helix overview and release tooling

Models

Falcon-H1-7B-FORGE-v2 Falcon-H1-7B-Instruct at 2.06 bpw / 3.48 GB. Runs on Apple Silicon in 3.73 GB resident at 75 tok/s decode. English only — this is an engineering artifact from the compression work, not a SEA-language model.

Read the model card before using it. It needs repeat_last_n = 2048; llama.cpp's default of 64 is too small to suppress the cross-turn repetition these models fall into, and the usual sampling triple without it reproduces the bug.

How we report numbers

Measurements come with their conditions attached, and with the cases that did not work. Things we published that contradicted our own assumptions:

Where a claim is not measured, we say so rather than rounding it up.

Contact

sebastian@sg-systems.co