Android — Google Play
Installed from Play like any other app, updates and all, on Android 8.0 or newer. Tap the button, install, done.
Your AI sidekick for everything you say out loud
Point it at a meeting, a conversation, or nothing in particular. SID3KICK records, transcribes, summarizes, pulls out the calendar events and to-dos you mentioned, and rolls the whole day into a digest you can actually search later. On your device by default — the offline engine never sends a byte anywhere.
7-day free trial · then $5/month · plus your own AI API usage · $0 if you run it fully offline
No invite, no waitlist, no email round-trip. Android installs straight from Google Play; iPhone and iPad from the App Store. Both are build 1.10.0 — the platforms release together and always carry the same version.
Installed from Play like any other app, updates and all, on Android 8.0 or newer. Tap the button, install, done.
A native app, not a wrapper — the same screens, every feature, the same themes and the same subscription, on iOS 16 or newer. Life Assist, camera, the map, sync and on-device AI all ship on iPhone too; the two things iOS has no API for (resuming after a reboot, a floating bubble over other apps) are the only Android extras.
What you get: every feature, no caps — a $5/month subscription with a 7-day free trial, priced identically on both stores. Nothing here is a time-limited demo. Found something broken? Tell us — it goes straight to the person who fixes it.
Tap to record when you know you want a transcript. Or leave it running and let it work out for itself when someone's actually talking.
Start recording and the pipeline takes over — transcribe, summarize, title, and scan for anything that sounds like a date or an appointment. On Android an optional floating bubble sits over other apps so you never have to open SID3KICK to hit record.
It listens in the background with the screen off and the phone in your pocket. Voice activity detection slices continuous audio into real speech, so you're not paying to transcribe silence — and each chunk runs the same pipeline. On iPhone it runs until you force-quit the app; on Android it can also restart itself after a reboot.
Summaries, the daily digest, recall chat and photo description can run on the phone itself — nothing leaves the device, nothing is billed. Bring your own model: a Gemma 4 file on Android, and on iPhone either Apple's built-in model (iOS 26, iPhone 15 Pro or newer) or a Gemma 4 GGUF. Each task chooses on-device or cloud independently. Needs a recent high-end phone.
It transcribes what's actually being spoken, not what your phone's locale expects. Then pin one output language and every title, summary, and digest comes back in it — which also inoculates you against cloud models hallucinating a wrong-language transcript on noisy audio.
Dates, times, and appointments get pulled out of what you said and surfaced as editable one-tap calendar suggestions. They write straight to your device's own calendar — no Google Cloud project, no OAuth screen, no API key.
"What did I say I'd send Priya?" "When was that dentist thing?" The Ask tab answers across every transcript at once, and because each capture is tagged with when — and optionally where — it happened, "yesterday" resolves against your real history instead of being guessed at.
To-dos get pulled out of what you actually said, land on a tickable list, show up in the daily digest, and — if you paste a token — push one-way to your Google Tasks default list.
Location tagging is off, country-only, or precise — your call, and off is the default: a fresh install takes no fix, stores no coordinates and never asks for the permission until you choose otherwise. Country-only reads the last-known fix with no GPS wake-up, keeps the place name, and blurs the coordinates to roughly a kilometre before storing anything.
Keep transcripts, summaries, to-dos, and digests in step across your devices through your Dropbox, Google Drive, or FTP server — never through anything we run. Each device writes its own snapshot and merges the others.
Dark and Light are ordinary Material 3. Console is not: black panels, asymmetric corners, condensed uppercase type, and synthesized interaction beeps generated at runtime — no ripped audio, no licensing headache.
Every stage is either on your phone or on a provider you hold the key to. There is no middle tier operated by us.
16 kHz mono audio, manual or voice-activated. Snap photos or short video alongside it — with video, only the audio is ever extracted and sent; the file itself never leaves the device.
On Android, Vosk on-device across 14 downloadable language models; on iPhone, Apple's built-in speech recognition — zero network calls, zero cost either way. Or route to OpenAI's cloud transcription when you want the accuracy and don't mind the upload.
Title, summary, extracted events, action items, and optional speaker diarization — through OpenAI, Anthropic, a local Ollama or LM Studio box, any endpoint that speaks the OpenAI chat-completions schema, or a model running on the phone itself.
Day-by-day digests, a map of where you were, describe-and-recall search for half-remembered conversations, and a chat box that answers across the lot.
Straight off a device — no mockups, no retouching. Dark theme shown; Light and Console are one Settings tap away. Android first, then the same screens on iPhone.
Same six tabs, same data, same subscription — captured on an iPhone 17 Pro running 1.10.0. Life Assist, camera, sync and on-device AI are all here, labelled early beta while they earn their stripes on Apple hardware.
Plenty of apps say "privacy-first" and mean "we encrypt it on the way to our servers." SID3KICK has a configuration where there are no servers.
On-device speech-to-text plus either On-device AI on the phone or a local Ollama or LM Studio box for summaries, and nothing leaves your network — or your phone. Not a degraded fallback — a first-class setup.
When you do use the cloud, you paste your own API key and talk to that provider directly. We are not a proxy, we never see the traffic, and there is no account to create.
Sync runs through storage you already own. Raw audio, photos, and video stay on the device that captured them; optional cloud archive and auto-prune only ever delete audio that's already been backed up.
Every subscription starts with a 7-day free trial — full access, cancel anytime. Two separate bills, and we're upfront about both. The subscription is ours. The AI usage is billed by whichever provider you point the app at — or it's nothing at all, because you're running the models on your own hardware.
The app
$5 / month
Starts with a 7-day free trial — try everything, cancel anytime before it ends and pay nothing.
One subscription, every feature, both platforms — the same $5/month and the same 7-day trial on Google Play and on the App Store.
Your AI usage
$0 + / month, at cost
Paid straight to your provider, never marked up by us. Run everything on-device and this line is genuinely $0.
| Setup | Transcription | Summaries | Per hour |
|---|---|---|---|
| Fully offline | On-device | On-device AI, or Ollama / LM Studio | $0.00 |
| Hybrid — the sweet spot | Vosk, on-device | Claude Haiku 4.5 | ~$0.02 |
| Cheapest all-cloud | gpt-4o-mini-transcribe | GPT-5.4 nano | ~$0.18 |
| Accuracy-first | whisper-1 | Claude Sonnet 5 | ~$0.41 |
How these are worked out: one hour of conversational speech is roughly
9,000 words, or about 12,000 tokens of transcript, plus around 600 tokens of generated
title, summary, and extracted events. Transcription is billed per audio minute
(gpt-4o-mini-transcribe $0.003/min, whisper-1 $0.006/min);
summaries are billed per token at each provider's published rate. These are
estimates, not quotes — your real numbers depend on how densely people
actually talk and on provider pricing at the time, which is exactly why the in-app price
table is editable.
| How you'd use it | Setup | AI usage | Total / month |
|---|---|---|---|
| Any amount, fully offline | On-device transcription + on-device or local model | $0.00 | $5.00 |
| ~3 h/day of talk, hybrid | Vosk + Claude Haiku 4.5 | ~$1.50 | ~$6.50 |
| ~1 h/day of meetings | Cloud transcription + GPT-5.4 nano | ~$5.50 | ~$10.50 |
| Passive all day, ~2 h of actual speech | Cloud transcription + GPT-5.4 nano | ~$11.00 | ~$16.00 |
| Heavy — ~3 h/day, all cloud | Cloud transcription + GPT-5.4 nano | ~$16.50 | ~$21.50 |
Notice the pattern: transcription is the expensive part, and it's the part you can move on-device for free. That's why the hybrid row is a rounding error — Vosk does the heavy lifting locally and a small hosted model handles the writing. Note too that passive mode bills for detected speech, not for hours of runtime: voice activity detection means eight hours in your pocket is nothing like eight hours of billed audio.
The app is live on both platforms — anyone can install it right now from Google Play or the App Store, same version, same features, released together. What gets this to a lot more people, a lot sooner, is the right people alongside it.
Integrations, distribution, and bundles. If you build for people who live in meetings, field work, journalism, accessibility, research, or healthcare admin — there's an obvious fit here, and an app deliberately built to talk to whatever endpoint you point it at.
Productivity, privacy, self-hosting, local-LLM, Android power-user, or second-brain audiences — this demos well and it holds up to scrutiny, because the offline mode is real and the cost tracker is honest. Early builds, direct access to the developer, and affiliate terms available.
A shipping product with a clear $5/month subscription live on both app stores, and no inference costs on our balance sheet — users bring their own keys or run the models on their own hardware. Happy to walk through the roadmap, the architecture, and the numbers.
Why the economics are unusual: because SID3KICK is on-device-first and bring-your-own-key, we don't pay for inference and we don't operate a storage backend. There's no per-user cloud cost quietly scaling underneath the subscription — which means the $5/month is margin, not a loss-leader waiting to be repriced.
The honest answer to "what do I need?": nothing but the app, if you run it on-device. An API key only enters the picture when you want a cloud model writing your summaries.
Install from Google Play or the App Store, run through onboarding (it asks for microphone permission and your recording consent), and you're on the Home screen. Tap the big button and you're recording. Then make one choice in Settings → Transcription: On-device — free and offline; Android downloads a Vosk language model the first time, iPhone uses Apple's built-in recognizer with nothing to download — or OpenAI (cloud), which needs an API key. Summaries, titles, events, and to-dos come from the AI provider you set up next (see below) — recording and transcription work fine before you do.
sk-.One note: Anthropic has no speech-to-text API, so transcription stays on-device (or on OpenAI) — Claude handles the summaries.
Yes — that's the whole point of the offline path. Transcription:
Settings → Transcription → On-device (Vosk) and download your
language. Summaries: run a local model with
Ollama or
LM Studio on a computer at home,
then in Settings → AI provider pick Ollama (or LM Studio) and
point the base URL at that machine — for Ollama that's
http://your-computer:11434. Nothing leaves your network, and the
usage bill is genuinely $0.
Yes — On-device AI, in early beta, under
Settings → AI provider. SID3KICK never downloads model weights
for you; the app links you to the model and you import the file yourself.
On Android that's a Gemma 4 .litertlm from Google's LiteRT
community page (about 2.6 GB). On iPhone, Apple's built-in model is used when
the phone has it (iOS 26, iPhone 15 Pro or newer) and needs no download at all;
otherwise import a Gemma 4 GGUF. Summaries, the daily digest, recall chat and
photo description each pick on-device or cloud separately, and the app states
the trade-off next to every switch. It wants a recent high-end phone with real
memory headroom — 8 GB is the comfortable floor.
The hybrid: Vosk on-device for transcription + a small cloud model for summaries. Transcription is the expensive part of every cloud setup, and it's the part the phone can do locally for free — so the hybrid lands around two cents per hour of speech. The pricing section has the full comparison, and you can switch setups any time without losing anything.
On the phone. Audio never leaves the device unless you chose cloud transcription; API keys are stored on-device and used to talk to your provider directly — there is no account with us and no server of ours in the middle. The privacy section and the full privacy policy spell it out.
Partnership, coverage, investment, or feedback — it all comes to the same inbox and it all gets a reply. (Just want the app? It's right here — no email needed.)
Prefer plain email? [email protected]