Skip to content

Deploy on-device
AI to your
mobile app in 1 hour

curl -fsSL https://raw.githubusercontent.com/zetic-ai/melange-cli/main/script/install.sh | sh

Or paste this command into your coding agent.

Your model.
On the device.

Melange automates model optimization, benchmarks on real devices, and generates SDK code to bring AI into your mobile app.

Built by engineers from

QualcommMicrosoftAmazonStanford UniversityUC Berkeley
Cornell UniversityEPFLSeoul National UniversityKAIST

Why On-Device AI

On-Device AI isn’t just fast.
It’s Profitable

No Latency

AI responds the instant the user acts, with no server round-trip. Real-time camera, voice, and live features feel snappy, and snappy keeps people coming back.

Full Privacy

Every inference runs on the user's device, so their data never leaves it. Privacy becomes a feature you can market and a box you can check for HIPAA and GDPR.

No server cost

No GPU servers to rent and no per-token cloud bills. Your cost stays flat as you scale, so margins hold instead of eroding with every new user.

Offline Access

The model keeps running with no signal, on a plane or in the field. You reach users bad connectivity would cut off, and nothing breaks when the network does.

Problem

Every device is different

Fragmented hardware, inconsistent performance, months of manual NPU tuning per chip. That's what building on-device AI looks like today. This is what Zetic solves.

Solution

From model to on-device AI.
Every step, automated.

Cutting 6+ months of engineering into 1 hour

Hardware-Aware Optimization in Melange — LFM2.5-350M

In practice

From the lab to local hardware.

Without Melange, we wouldn't have been able to demonstrate our model's feasibility in real-world situations running on local hardware.
William WangWilliam WangSpectra team · Stanford

How it works

3 Simple Steps to Deploy

Go from a raw AI Model to Shippable Mobile App Integration in 1 hour

01

Select or Upload
your own model

Upload your model, use a Hugging Face link, or choose from the model library.

02

Benchmark
across 200+ mobile devices

Compare measured performance on real devices and choose your deployment settings.

03

Deploy
by copying our SDK code block

Copy the generated SDK code into your app and run your model on-device.

Core Capabilities

Built into Melange

Dynamic Routing

The SDK identifies the device and downloads the best-performing runtime, data type, and processor combination for your model.

Automated Benchmarking

Real accuracy and latency reports across 200+ physical devices.

Multi-Runtime Acceleration

CPU, GPU, and NPU hybrid support, handled automatically.

Automated Optimization

Model conversion, quantization, and optimization in one pipeline.

FAQ

Frequently Asked Questions

Get answers to common questions here

Do I need to retrain my model to use Melange?

No. We support TorchScript, TensorFlow and ONNX models directly. Our platform automatically handles conversion and quantization for on-device execution without needing your training data or altering weights

Why use Melange instead of free open-source tools like TFLite or CoreML?

Open-source tools force you to manage fragmented pipelines: wrestling with CoreML for iOS, TFLite/LiteRT for Android, and endless hardware variations across devices. Melange unifies this into a single, streamlined workflow. We replace months of manual engineering with automated, deep NPU optimization that generic tools miss, and solve hardware fragmentation by providing real-world benchmarking across 200+ physical devices.

How much cost savings can be achieved by using Melange?

The savings are substantial because on-device AI decouples your user growth from your server costs. With cloud GPUs, doubling your users means doubling your infrastructure bill. By shifting computation to the user's device, your ongoing inference cost per new user drops to near zero. You stop paying indefinite cloud rent and shift to a highly scalable, predictable cost model.

Is on-device AI actually faster than a powerful cloud GPU server?

Often, yes, because it eliminates network latency. While cloud GPUs have more raw power, the round-trip data transfer slows them down. On-device NPU execution is instantaneous, providing a smoother experience regardless of internet connection.

What happens if a user’s phone is old and doesn't have an NPU?

Your app will still work. Our runtime engine automatically detects available hardware. It prioritizes the NPU for speed, but seamlessly falls back to the GPU or CPU on older devices to guarantee broad compatibility.

How is Melange different from Google AI Edge?

With Google AI Edge, you can upload a model, optimize it, and see benchmarks, but taking it to a shipped app means building the deployment yourself. Melange does it all in one platform: upload, optimize, benchmark on real devices, and deploy a production-ready SDK to iOS or Android, in minutes. One workflow, model to shipped app.

Get started

Two ways to get started

Start free on mobile. Scale to enterprise when you need more.

Start building with Melange

Upload your model or explore the library. Start with the dashboard or CLI.

Try Melange

Melange Enterprise

Have a custom deployment?

Custom DevTools help enterprise teams adapt on-device AI deployment to their own hardware, runtime, and product requirements.

Talk to our team