This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
Memories is a photo and video viewer for our Android TV. I built it for especially my mom who loves to browse through old photos and videos on TV 🙂
Just last week my mom asked if there's an app that can analyze a picture and tell me where it was taken. I thought that would be a fun weekend project to try out Gemma for visual intelligence. Now, this app has that capability plus a lot more!
You plug the photo SSD into the TV and the app opens by itself. Everything works from the couch with the regular remote. You can browse folders and run slideshows. It even plays the 4K Dolby Vision videos from our trips. Every photo also shows the date and the place it was taken. One of the other reasons why I started working on this app was the sheer lack of a good photo viewer for Android TV. The built-in one is slow and clunky. I wanted something that was fast, simple, and fun to use.
This weekend I added the part I'm most excited about. On any photo you can press Down on the remote and just ask about it. Things like "What is this place called?" or "What's that tower on the left called?". You can type the question or hold the mic button and say it.
If you zoom into part of the photo first then the question is about that part. The answer comes from Gemma running on a Mac in the same house. It shows up in a couple of seconds and the TV reads it out loud.
Demo
Code
sbis04
/
memories_gemma
Media viewer for Android TV
Memories
A photo and video viewer for Android TV that lets you ask questions about your photos. The answers come from Gemma running on a computer in your home, so your photos never leave the house.
Features
- Made for the remote. Plug in your photo drive and the app opens by itself. Browse folders as a grid or list and pin the ones you use most.
- Slideshows with fade, Ken Burns or slide transitions. They resume where you left off.
- 4K, HDR and Dolby Vision video through the TV's own decoder.
- Date and place on every photo, read from its GPS data.
- Ask Gemma. Press Down on a photo and ask about it by typing or with the remote's mic. Zoom in first to ask about one part. The answer streams in and the TV reads it aloud.
Memories.Gemma.Demo.mov
How I Built It
The app is written in Flutter and runs on a Sony Bravia with Android TV 14. There's no touch screen and no mouse. Every screen had to work with just the D-pad.
For the AI part I'm using Gemma 4 (the E2B size) through Ollama. It runs on my MacBook and the TV talks to it over the home Wi-Fi:
TV app ── photo + question over LAN ──▶ Mac: Ollama → gemma4 ▲ │
└──────── answer, streamed word by word ◀───────────────┘
then read aloud by the TV's built-in text-to-speech
Here's roughly what happens when you ask something:
- The app sends Gemma the photo shrunk to about 1024px. If you're zoomed in it also sends a sharper crop of what's on screen. Gemma is told to focus on that crop. That's what makes "what's that building?" work on a wide skyline shot.
- The first question also carries what the app already knows about the photo. That's the date and the place name from the photo's GPS and the album folder. So for our Shanghai photo Gemma came back with "This photo was taken in Shanghai, China, on September 14, 2026… landmarks like the Oriental Pearl Tower" instead of guessing.
- Follow-up questions keep the earlier ones as context. So "and the one next to it?" works.
- The answer streams in over Ollama's
/api/chatas Gemma writes it. Once it's done the TV reads it out with Android's built-in voice.
Speed took the most work. My first version took anywhere from 10 to 30 seconds per answer. On a TV that feels like the app has frozen. When I looked at where the time went I found Gemma 4 spending around 500 tokens "thinking" before answering even "what is this?". I turned thinking off (think: false) and switched to the smaller E2B model. For photo questions I honestly couldn't tell the difference in the answers:
| Setup (M3 Max, model already loaded) | Full answer | First words on screen |
|---|---|---|
gemma4 (E4B), thinking on |
10 to 29 s | nothing until the whole answer is done |
gemma4 (E4B), thinking off |
about 5 s | under 1 s (streamed) |
gemma4:e2b, thinking off |
about 2.5 to 3 s | under 1 s |
I also ask Ollama to keep the model in memory for 30 minutes. That way only the first question of the evening has to wait for it to load.
If you want to try it then the Mac side is just this:
brew install ollama
ollama pull gemma4:e2b
OLLAMA_HOST=0.0.0.0 ollama serve # so the TV can reach it over Wi-Fi
Then put the Mac's IP address in Memories under Settings > Gemma server.
Why Does Open Innovation Matter?
The first version of this feature actually used a cloud API. Every question sent the photo off to a hosted model. You also needed an API key to use it. Switching to Gemma changed a few things that matter a lot for an app like this.
The big one is privacy. These are our family photos. They have kids and faces in them along with our house and every place we've lived. Now when you ask about a photo it only travels from the TV to the Mac in the next room. Nothing gets uploaded and I don't have to read anyone's data policy.
It's also free to run. Nobody has to type an API key with a TV remote. There's no bill or rate limit to think about either.My parents will never need to know what a token is 😂
It won't change on me. The cloud model I started with sat behind a "latest" alias. That's nice right up until it starts answering differently one day. The Gemma weights are sitting on my disk. If it works today it should keep working.
Running it locally also meant I could dig into it. I found the thinking problem myself and measured every change I made. Picking the smaller model and adding streaming came out of that. Trying another model is just a setting in the app. I use gemma4:e2b for speed. gemma4 gives slightly nicer wording and gemma4:26b is there if you have a beefier machine.
The catch is that a computer at home has to be switched on. We mostly use the TV in the evening and the Mac is right there anyway. I'm happy with that if it means our photos stay at home.
My Agent Session
I did the Gemma work in one afternoon with Claude Code. I swapped the cloud API for Ollama and set up Gemma on the Mac. Then I chased down the slow answers and tested everything on the actual TV. I even had it write the little script I used to record the demo.
can we add gemma to this project to make it eligible for this hackathon: https://dev.to/challenges/hacktoberfest-weekend-2026-10-01
[Image #1] [Image #2] [Image #3] [Image #4] [Image #5]
[screenshot]
[screenshot]
[screenshot]
[screenshot]
[screenshot]
$ git log --oneline | head -20 && ls && cat pubspec.yaml && find lib -name '*.dart' | xargs wc -l | sort -n | tail -40 && ls *.md 2>/dev/null; cat CLAUDE.md 2>/dev/null | head -100
4bfe9c4 Read Gemini's answers aloud with the TV's built-in text-to-speech
82e7ea5 Ask Gemini about photos; zoom in paused slideshows; fix double Back
bed29a4 Open Memories when the photo SSD is plugged in
baedb5f Show time after the date on loading placeholders
8e4b916 Fix 4K/Dolby Vision video stutter: free decoders early, native caption
daffbc8 Smoother video playback; save slideshow resume point on every slide
53bb028 Folder actions menu (pin/rename/delete) and date-sorted folders
5698fc3 Persist folder summaries so counts and covers show instantly
94d6792 Read drives directly on Android; list view, pins, priority loading
31d727c set to latest stable flutter
ed0c773 Instant folder browsing on Android + folder drawer, zoom and Back fixes
2835821 Viewer: aspect-fit native video, media-location grant, subtler caption
15c7f0e Viewer/slideshow: date + reverse-geocoded location caption
39c48af Slideshow: 95%-quality images + hold current slide until next is decoded
bbde43a Viewer: focusable video controls + wakelock while playing
c2b6c77 Native ExoPlayer video (HDR/DV), faster thumbnails, focus & nav fixes
8567f5c Viewer: add 200% zoom step
7bd5a4d Viewer: stepped zoom + D-pad pan
758394d Polish: faster thumbnails, slideshow pre-warm/resume, grid fill, fixes
9dbae65 Auto-refresh as Android indexes the USB drive
README.md
analysis_options.yaml
android
assets
build
lib
macos
pubspec.lock
pubspec.yaml
test
tv_gallery.iml
name: tv_gallery
description: "A new Flutter project."
# The
…[7393 more chars]
$ cat lib/services/gemini_service.dart README.md && grep -n -i "gemini\|apikey\|api_key" -r lib --include=*.dart | grep -v "^lib/services/gemini_service.dart" && git show --stat HEAD~1 | head -30
Exit code 1
import 'dart:convert';
import 'dart:io';
import 'dart:math' as math;
import 'dart:ui' as ui;
import 'package:flutter/painting.dart';
import 'package:path/path.dart' as p;
import '../models/media_entry.dart';
import '../widgets/media_caption.dart';
import 'location_service.dart';
import 'settings_controller.dart';
import 'thumbnail_service.dart';
/// A conversation with Gemini about one photo (or a video's poster frame).
///
/// Uses the Gemini API directly over HTTPS with the user's API key from
/// Settings. The image is sent downscaled (~1024px) along with what the app
/// already knows — date, place, folder — so questions like "where was this?"
/// get grounded answers. Follow-ups keep the history, so "and that building
/// on the left?" works.
class GeminiChat {
GeminiChat(this.entry, {Rect? focus})
: focus = focus == null || focus == const Rect.fromLTRB(0, 0, 1, 1)
? null
: focus;
final MediaEntry entry;
/// The zoomed-in region (normalized 0–1), or null for the whole photo.
final Rect? focus;
bool get zoomed => focus != null;
/// Newest Flash model (Google keeps this alias pointing at it).
static const model = 'gemini-flash-latest';
final List<Map<String, Object>> _history = [];
String? _imageData;
String? _cropData;
Future<void>? _prepared;
/// Prepares the image(s) to send — called as soon as the Ask screen opens,
/// so they're ready by the time a question has been typed or spoken.
Future
…[8049 more chars]
$ grep -rn -i "gemini" lib | grep -v "^lib/services/gemini_service.dart"; grep -n "minSdk\|targetSdk" android/app/build.gradle* ; cat android/app/src/main/AndroidManifest.xml | head -40
lib/screens/settings_screen.dart:179: _SectionHeader('Gemini'),
lib/screens/settings_screen.dart:181: title: 'Gemini API key',
lib/screens/settings_screen.dart:182: subtitle: s.geminiApiKey == null
lib/screens/settings_screen.dart:185: : 'Set (…${_tail(s.geminiApiKey!)}) — stored only '
lib/screens/settings_screen.dart:189: label: s.geminiApiKey == null ? 'Add' : 'Change',
lib/screens/settings_screen.dart:190: onPressed: () => _editGeminiKey(context),
lib/screens/settings_screen.dart:195: subtitle: 'Speak Gemini’s answers with the TV voice',
lib/screens/settings_screen.dart:228: Future<void> _editGeminiKey(BuildContext context) async {
lib/screens/settings_screen.dart:232: title: 'Gemini API key',
lib/screens/settings_screen.dart:233: initial: settings.geminiApiKey ?? '',
lib/screens/settings_screen.dart:237: if (key != null) await settings.setGeminiApiKey(key);
lib/screens/settings_screen.dart:468: static const _sample = 'Hi! This is how Gemini’s answers will sound.';
lib/screens/slideshow_screen.dart:314: /// Pauses (if needed) and asks Gemini about the current slide — the
lib/screens/slideshow_screen.dart:328: /// arrows pan while zoomed. Up goes to the top-bar chips, Down asks Gemini.
lib/screens/viewer_screen.dart:240: // Down asks Gemini (zoomed: pan down first, ask at the bottom ed
…[5245 more chars]
$ which ollama llama-server lms 2>/dev/null; ollama list 2>/dev/null; sysctl -n hw.memsize machdep.cpu.brand_string; flutter --version 2>/dev/null | head -1
ollama not found
llama-server not found
lms not found
38654705664
Apple M3 Max
Flutter 3.47.6 • channel stable • https://github.com/flutter/flutter.git
{
"questions": [
{
"question": "Right now \"Ask Gemini\" sends each photo to Google's Gemini API. Where should Gemma run instead? (This decides how strong the \"why open matters\" part of your post is.)",
"header": "Gemma host",
"multiSelect": false,
"options": [
{
"label": "Home server via Ollama (Recommended)",
"description": "Gemma runs on your Mac (or any machine on the home Wi-Fi) through Ollama. The TV sends photos over the LAN, so family photos never leave the house, it's free, and you can swap models. Needs the Mac on while you use the feature. The app also gets a 'Gemma server address' setting.",
"preview": "TV (Memories app)\n │ photo + question, LAN only\n ▼\nMac: ollama serve → gemma3 (vision)\n │\n ▼\nanswer → read aloud on TV\n\nNo cloud, no API key."
},
{
"label": "On the TV itself",
"description": "Gemma 3n runs on-device via flutter_gemma/LiteRT. Fully offline, but the Bravia's MediaTek chip and RAM may make it very slow or crash. Needs a ~3GB model download onto the TV. Highest risk to finish this weekend.",
"preview": "TV (Memories app)\n └─ flutter_gemma → Gemma 3n E2B (on-device)\n\n~3GB model on the TV, slow on TV SoC,\nmay not fit in memory."
},
{
"label": "Gemma via Gemini API",
"description": "Same API key, just change the model to a hosted Gemma. About 20 lines of code, but photos still go to Google, so it's a weak fit for 'open AI at its core'.",
"preview": "TV → generativelanguage.googleapis.com\n model: gemma-3-27b-it\n\nSame key, photos still go to the cloud."
},
{
"label": "Ollama + keep Gemini option",
"description": "Same as the first option, but Settings has a Gemma (local) / Gemini (cloud) toggle. Gemma is the default and Gemini stays as a fallback when the Mac is off."
}
]
}
]
}
Your questions have been answered: "Right now "Ask Gemini" sends each photo to Google's Gemini API. Where should Gemma run instead? (This decides how strong the "why open matters" part of your post is.)"="Home server via Ollama (Recommended)" selected preview:
TV (Memories app)
│ photo + question, LAN only
▼
Mac: ollama serve → gemma3 (vision)
│
▼
answer → read aloud on TV
No cloud, no API key.. You can now continue with these answers in mind.
{
"query": "select:WebSearch",
"max_results": 1
}
[image]
{
"query": "ollama gemma latest vision model tag 2026",
"mode": "standard"
}
Web search results for query: "ollama gemma latest vision model tag 2026"
Links: [{"title":"Best Ollama Models (2026): 25+ Ranked by VRAM, Context, and SWE-Bench (August 2026)","url":"https://www.morphllm.com/best-ollama-models"},{"title":"How to Run Gemma 4 with Ollama: Step-by-Step Setup Guide (2026)","url":"https://codersera.com/blog/how-to-run-gemma-4-with-ollama-setup-guide/amp/"},{"title":"top open source models ollama","url":"https://www.promptquorum.com/local-llms/top-open-source-models-ollama"},{"title":"which ollama models support vision","url":"https://www.promptquorum.com/es/prompt-bites/which-ollama-models-support-vision"},{"title":"www.serverman.co.uk","url":"https://www.serverman.co.uk/?p=4092"},{"title":"which ollama models support vision","url":"https://www.promptquorum.com/pt/prompt-bites/which-ollama-models-support-vision"},{"title":"top open source models ollama","url":"https://www.promptquorum.com/es/local-llms/top-open-source-models-ollama"},{"title":"Ir al contenido principal","url":"https://www.javadex.es/blog/ollama-modelos-ranking-completo-mejores-modelos-local-2026"},{"title":"which ollama models support vision","url":"https://www.promptquorum.com/prompt-bites/which-ollama-models-support-vision"}]
Based on the search results, here's what I found regarding Ollama Gemma vision models in 2026:
## Gemma 4 Vision Model
Google DeepMind released Gemma 4 on April 2, 2026, and it comes in four sizes — E2B, E4B, E12B, and E27B — all natively multimodal, m
…[585 more chars]
$ sed -n 30,130p lib/services/settings_controller.dart && sed -n 170,185p lib/services/settings_controller.dart && sed -n 150,260p lib/screens/settings_screen.dart && sed -n 440,480p lib/screens/settings_screen.dart
static const _kSlideShuffle = 'slideshowShuffle';
static const _kSource = 'sourcePath';
static const _kSourceLabel = 'sourceLabel';
static const _kLastFolder = 'lastFolderPath';
static const _kSlideLastFolder = 'slideshowLastFolder';
static const _kSlideLastItem = 'slideshowLastItem';
static const _kFolderCounts = 'folderMediaCounts';
static const _kListView = 'listView';
static const _kPinned = 'pinnedFolders';
static const _kGeminiKey = 'geminiApiKey';
static const _kReadAloud = 'readAnswersAloud';
static const _kAnswerVoice = 'answerVoice'; // "name|locale"
// ---- State (defaults) ----
ThemeMode _themeMode = ThemeMode.dark;
int _gridZoom = 2; // 0..maxZoom (2 = default, one step below "Medium")
// Media defaults to chronological order, oldest first.
SortBy _sortBy = SortBy.date;
bool _sortDesc = false;
bool _showTitles = false;
bool _showCaption = true;
bool _autoplayVideos = true;
bool _muteVideo = false;
int _slideshowSeconds = 6;
SlideshowTransition _slideshowTransition = SlideshowTransition.kenBurns;
bool _slideshowLoop = true;
bool _slideshowShuffle = false;
String? _sourcePath;
String? _sourceLabel;
String? _lastFolderPath;
String? _slideshowLastFolder;
String? _slideshowLastItem;
Map<String, int> _folderCounts = {};
bool _listView = false;
List<String> _pinned = [];
String? _geminiApiKey;
bool _readAloud = true;
String? _answerVoice;
// ---- Getters ----
ThemeMode get themeMode => _themeM
…[8459 more chars]
$ cat lib/widgets/ask_panel.dart; sed -n 20,40p lib/services/answer_voice.dart; cat macos/Runner/*.entitlements; ls android/app/src/main/res/xml 2>/dev/null; grep -n "Cleartext\|networkSecurity" -r android/app/src
import 'dart:async';
import 'dart:io';
import 'dart:ui' as ui;
import 'package:flutter/material.dart';
import 'package:flutter/services.dart';
import 'package:flutter_animate/flutter_animate.dart';
import 'package:lucide_icons_flutter/lucide_icons.dart';
import '../core/app_theme.dart';
import '../core/transitions.dart';
import '../models/media_entry.dart';
import '../screens/settings_screen.dart';
import '../services/answer_voice.dart';
import '../services/gemini_service.dart';
import '../services/settings_controller.dart';
import '../services/thumbnail_service.dart';
import 'focusable.dart';
/// Opens "Ask Gemini" full screen over the blurred photo: a centred question
/// field (focused, so the TV keyboard — with its voice button — comes straight
/// up), a preview of exactly what Gemini is shown (the zoomed-in part, if
/// zoomed), quick suggestions and the conversation. Back closes it.
/// [focus] is the zoomed-in region (normalized).
Future<void> showAskPanel(
BuildContext context,
MediaEntry entry, {
Rect? focus,
}) {
return showGeneralDialog<void>(
context: context,
barrierDismissible: false,
barrierLabel: 'Ask Gemini',
barrierColor: Colors.transparent,
transitionDuration: AppTheme.focusAnim * 2,
pageBuilder: (context, _, _) => _AskPanel(entry, focus: focus),
transitionBuilder: (context, animation, _, child) =>
FadeTransition(opacity: animation, child: child),
);
}
class _AskPanel extends StatefulWidget {
const _AskPa
…[20454 more chars]
$ grep -rn "allowSlash" lib | head; grep -n "Future<String?> showTextInputDialog" -A40 -r lib | head -60; sed -n 1,30p lib/screens/settings_screen.dart
lib/screens/settings_screen.dart:235: allowSlash: true,
lib/widgets/folder_actions.dart:77: bool allowSlash = false,
lib/widgets/folder_actions.dart:86: allowSlash: allowSlash,
lib/widgets/folder_actions.dart:285: required this.allowSlash,
lib/widgets/folder_actions.dart:292: final bool allowSlash;
lib/widgets/folder_actions.dart:331: if (!widget.allowSlash) FilteringTextInputFormatter.deny(RegExp('[/]')),
lib/widgets/folder_actions.dart:71:Future<String?> showTextInputDialog(
lib/widgets/folder_actions.dart-72- BuildContext context, {
lib/widgets/folder_actions.dart-73- required String title,
lib/widgets/folder_actions.dart-74- String initial = '',
lib/widgets/folder_actions.dart-75- String? hint,
lib/widgets/folder_actions.dart-76- String confirmLabel = 'Save',
lib/widgets/folder_actions.dart-77- bool allowSlash = false,
lib/widgets/folder_actions.dart-78-}) {
lib/widgets/folder_actions.dart-79- return _showPanel<String>(
lib/widgets/folder_actions.dart-80- context,
lib/widgets/folder_actions.dart-81- builder: (context) => _TextForm(
lib/widgets/folder_actions.dart-82- title: title,
lib/widgets/folder_actions.dart-83- initial: initial,
lib/widgets/folder_actions.dart-84- hint: hint,
lib/widgets/folder_actions.dart-85- confirmLabel: confirmLabel,
lib/widgets/folder_actions.dart-86- allowSlash: allowSlash,
lib/widgets/folder_actions.dart-87- ),
lib/widgets/folder_actions.dart-88- );
lib/widgets/folder_actions
…[2401 more chars]
$ git mv lib/services/gemini_service.dart lib/services/gemma_service.dart
(Bash completed with no output)
Going with Gemma on a home Mac through Ollama, as you picked. I'm now swapping the app's Gemini code over to call it.
actually first commti the inital app before making these new cahnges, create a commit called "Add Memories TV app"
$ git mv lib/services/gemma_service.dart lib/services/gemini_service.dart && git status --short && git remote -v && git rev-list --count HEAD
Exit code 128
fatal: not under version control, source=lib/services/gemma_service.dart, destination=lib/services/gemini_service.dart
$ git status; ls lib/services/ | grep gem; git remote -v; git log --oneline | head -3
On branch main
No commits yet
Untracked files:
(use "git add <file>..." to include in what will be committed)
.fvmrc
.gitignore
.metadata
.vscode/
README.md
analysis_options.yaml
android/
assets/
lib/
macos/
pubspec.lock
pubspec.yaml
test/
nothing added to commit but untracked files present (use "git add" to track)
gemma_service.dart
fatal: your current branch 'main' does not have any commits yet
The git history has been reset (no commits yet). First I'll undo my file rename so the commit has the original app, then check that nothing secret or generated gets staged.
$ mv lib/services/gemma_service.dart lib/services/gemini_service.dart && cat .gitignore | head -50 && git add -A && git status --short | wc -l && git status --short | grep -v "^A \(lib\|android\|macos\|assets\|test\)/" ; grep -rnE "AIza[0-9A-Za-z_-]{20,}" $(git diff --cached --name-only) 2>/dev/null | head
# Miscellaneous
*.class
*.log
*.pyc
*.swp
.DS_Store
.atom/
.build/
.buildlog/
.history
.svn/
.swiftpm/
migrate_working_dir/
# IntelliJ related
*.iml
*.ipr
*.iws
.idea/
# The .vscode folder contains launch configuration and tasks you configure in
# VS Code which you may wish to be included in version control, so this line
# is commented out by default.
#.vscode/
# Flutter/Dart/Pub related
**/doc/api/
**/ios/Flutter/.last_build_id
.dart_tool/
.flutter-plugins-dependencies
.pub-cache/
.pub/
/build/
/coverage/
# Symbolication related
app.*.symbols
# Obfuscation related
app.*.map.json
# Android Studio will place build artifacts here
/android/app/debug
/android/app/profile
/android/app/release
# FVM Version Cache
.fvm/
# Golden test failure diffs
test/failures/
134
A .fvmrc
A .gitignore
A .metadata
A .vscode/settings.json
A README.md
A analysis_options.yaml
A pubspec.lock
A pubspec.yaml
$ git diff --cached --name-only | grep -iE "local.properties|key.properties|\.jks|\.keystore|google-services|\.env" ; git commit -q -m "Add Memories TV app
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>" && git log --stat --oneline | tail -1
134 files changed, 15268 insertions(+)
$ git mv lib/services/gemini_service.dart lib/services/gemma_service.dart
(Bash completed with no output)
$ cat > <scratch>/head.dart <<'EOF'
/// A conversation with Gemma about one photo (or a video's poster frame).
///
/// Gemma, Google's open-weight model, runs on a computer at home through
/// Ollama (Settings → Gemma server); the TV talks to it over the local
/// network, so photos never leave the house and there's no API key or bill.
/// The image is sent downscaled (~1024px) along with what the app already
/// knows — date, place, folder — so questions like "where was this?" get
/// grounded answers. Follow-ups keep the history, so "and that building on
/// the left?" works.
class GemmaChat {
GemmaChat(this.entry, {Rect? focus})
: focus = focus == null || focus == const Rect.fromLTRB(0, 0, 1, 1)
? null
: focus;
final MediaEntry entry;
/// The zoomed-in region (normalized 0–1), or null for the whole photo.
final Rect? focus;
bool get zoomed => focus != null;
final List<Map<String, Object>> _history = [];
String? _imageData;
String? _cropData;
Future<void>? _prepared;
/// Prepares the image(s) to send — called as soon as the Ask screen opens,
/// so they're ready by the time a question has been typed or spoken.
Future<void> prepare() => _prepared ??= () async {
final crop = focus;
final results = await Future.wait([
_image(),
crop == null ? Future<String?>.value() : _crop(crop),
]);
_cropData = results[1];
}();
static bool get hasServer => SettingsController.instance.gemmaServer != null;
/// Asks [question] about the photo; returns the answer text. Throws a
/// [GemmaException] with a user-facing message on failure.
Future<String> ask(String question) async {
final settings = SettingsController.instance;
final server = settings.gemmaServer;
if (server == null) {
throw const GemmaException('Add your Gemma server in Settings.');
}
final model = settings.gemmaModel;
final Map<String, Object> turn;
if (_history.isEmpty) {
await prepare();
final crop = _cropData;
turn = {
'role': 'user',
'content': [
await _context(),
if (crop != null)
'The viewer has zoomed in: the second image is the part of the '
'photo on screen. Focus on it; the first image is the whole '
'photo for context.',
question,
].join('\n\n'),
'images': [await _image(), ?crop],
};
} else {
turn = {'role': 'user', 'content': question};
}
final body = jsonEncode({
'model': model,
'messages': [
{
'role': 'system',
'content': 'You help someone look through their own photos and '
'videos on a TV. Answer about the attached picture in plain '
'text (no markdown), concisely: usually 2–5 sentences, readable '
'from a couch. Use the date, place and folder given as context '
'when relevant. If you ar
…[2753 more chars]
(Bash completed with no output)
Now the settings: replacing the API key with a server address and a model name.
$ python3 - <<'EOF'
p='lib/services/settings_controller.dart'
s=open(p).read()
def r(a,b):
global s
assert a in s, a
s=s.replace(a,b)
r(" static const _kGeminiKey = 'geminiApiKey';\n",
" static const _kGemmaServer = 'gemmaServer';\n static const _kGemmaModel = 'gemmaModel';\n")
r(" String? _geminiApiKey;\n", " String? _gemmaServer;\n String _gemmaModel = defaultGemmaModel;\n")
a=s.index(" /// Gemini API key (from aistudio.google.com)")
b=s.index(" /// Read Gemini's answers aloud")
s=s[:a]+''' /// The Ollama server running Gemma on the home network, as a base URL
/// (e.g. http://192.168.1.20:11434), or null if not set up yet.
String? get gemmaServer => _gemmaServer;
/// Accepts a bare host or IP ("192.168.1.20", "mac.local") and fills in
/// Ollama's scheme and port.
Future<void> setGemmaServer(String? address) async {
var a = address?.trim() ?? '';
while (a.endsWith('/')) {
a = a.substring(0, a.length - 1);
}
if (a.isNotEmpty) {
if (!a.contains('://')) a = 'http://$a';
final uri = Uri.tryParse(a);
if (uri != null && uri.host.isNotEmpty && !uri.hasPort) {
a = uri.replace(port: 11434).toString();
}
}
_gemmaServer = a.isEmpty ? null : a;
if (_gemmaServer == null) {
await _prefs.remove(_kGemmaServer);
} else {
await _prefs.setString(_kGemmaServer, _gemmaServer!);
}
notifyListeners();
}
/// Gemma 4's default (E4B) build: multimodal and quick on a laptop.
static const defaultGemmaModel = 'gemma4';
/// Ollama model tag to ask (e.g. gemma4, gemma4:26b).
String get gemmaModel => _gemmaModel;
Future<void> setGemmaModel(String? model) async {
final m = model?.trim() ?? '';
_gemmaModel = m.isEmpty ? defaultGemmaModel : m;
await _prefs.setString(_kGemmaModel, _gemmaModel);
notifyListeners();
}
'''+s[b:]
r(" /// Read Gemini's answers aloud"," /// Read Gemma's answers aloud")
r(" _geminiApiKey = _prefs.getString(_kGeminiKey);\n",
" _gemmaServer = _prefs.getString(_kGemmaServer);\n _gemmaModel = _prefs.getString(_kGemmaModel) ?? defaultGemmaModel;\n")
open(p,'w').write(s)
p='lib/screens/settings_screen.dart'
s=open(p).read()
a=s.index(" _SectionHeader('Gemini'),")
b=s.index(" _SettingTile(\n title: 'Read answers aloud',")
s=s[:a]+''' _SectionHeader('Gemma'),
_SettingTile(
title: 'Gemma server',
subtitle: s.gemmaServer == null
? 'A computer at home running Ollama — photos '
'stay on your network'
: s.gemmaServer!,
trailing: IconFocusButton(
icon: LucideIcons.server,
label: s.gemmaServer == null ? 'Add' : 'Change',
onPressed: () => _editGemmaServer(context)
…[2567 more chars]
android/app/src/main/AndroidManifest.xml:11: <!-- Network access for "Ask Gemini" (Flutter's template only grants it
android/app/src/main/AndroidManifest.xml:15: visibility) to read Gemini's answers aloud. -->
$ python3 - <<'EOF'
p='android/app/src/main/AndroidManifest.xml'
s=open(p).read()
s=s.replace(''' <!-- Network access for "Ask Gemini" (Flutter's template only grants it
in debug/profile builds, so release builds couldn't connect). -->''',''' <!-- Network access for "Ask Gemma" (Flutter's template only grants it
in debug/profile builds, so release builds couldn't connect). -->''')
s=s.replace("visibility) to read Gemini's answers aloud.","visibility) to read Gemma's answers aloud.")
old=''' android:banner="@drawable/tv_banner">'''
assert old in s
s=s.replace(old,''' android:banner="@drawable/tv_banner"
android:usesCleartextTraffic="true">''')
s=s.replace(''' <application
android:label="Memories"''',''' <!-- usesCleartextTraffic: Gemma runs on a computer on the home network
via Ollama, which speaks plain HTTP. -->
<application
android:label="Memories"''')
open(p,'w').write(s)
EOF
flutter analyze 2>&1 | tail -15
Upgrading analysis_options.yaml to exclude build and platform directories.
Analyzing tv_gallery_gemma...
No issues found! (ran in 8.4s)
$ git status --short && git diff analysis_options.yaml
M analysis_options.yaml
M android/app/src/main/AndroidManifest.xml
M lib/screens/settings_screen.dart
M lib/screens/slideshow_screen.dart
M lib/screens/viewer_screen.dart
M lib/services/answer_voice.dart
RM lib/services/gemini_service.dart -> lib/services/gemma_service.dart
M lib/services/settings_controller.dart
M lib/widgets/ask_panel.dart
M lib/widgets/zoomable_image.dart
diff --git a/analysis_options.yaml b/analysis_options.yaml
index 0d29021..13c3ff6 100644
--- a/analysis_options.yaml
+++ b/analysis_options.yaml
@@ -7,6 +7,11 @@
# The following line activates a set of recommended lints for Flutter apps,
# packages, and plugins designed to encourage good coding practices.
+analyzer:
+ exclude:
+ - build/**
+ - android/**
+ - macos/**
include: package:flutter_lints/flutter.yaml
linter:
$ git checkout analysis_options.yaml && git diff --stat && which brew
Updated 1 path from the index
android/app/src/main/AndroidManifest.xml | 9 +-
lib/screens/settings_screen.dart | 58 ++++++----
lib/screens/slideshow_screen.dart | 4 +-
lib/screens/viewer_screen.dart | 6 +-
lib/services/answer_voice.dart | 2 +-
lib/services/gemma_service.dart | 185 +++++++++++++------------------
lib/services/settings_controller.dart | 54 ++++++---
lib/widgets/ask_panel.dart | 40 +++----
lib/widgets/zoomable_image.dart | 2 +-
9 files changed, 192 insertions(+), 168 deletions(-)
/opt/homebrew/bin/brew
"Ask Gemini" is now "Ask Gemma". Gemma runs on a computer at home through Ollama, and the TV sends it the photo over the home Wi-Fi. flutter analyze passes, but I haven't sent it a real question yet because Ollama isn't installed on your Mac. These changes aren't committed; the commit before them is "Add Memories TV app".
What changed
- lib/services/gemma_service.dart (renamed from gemini_service.dart): it now calls Ollama's /api/chat instead of Google's API. It still sends the downscaled photo, the zoomed-in crop, and the date, place and folder, and follow-up questions still work.
- It asks Ollama to keep the model loaded for 30 minutes, so later questions don't wait for it to reload.
- The first question can take up to 3 minutes while the model loads.
- Error messages say what to fix: Mac off or Ollama not on the network, model not downloaded (with the exact ollama pull command), or the model being too slow.
- Settings: the Gemini API key is replaced by "Gemma server" and "Model".
- For the server you can type just an IP or name like [REDACTED] and it fills in http://…:11434.
- The model defaults to gemma4, Gemma 4's smaller default size (~9.6 GB).
- Text on screen: every "Gemini" is now "Gemma". The "no server set" screen explains the setup in one sentence.
- AndroidManifest.xml: I turned on plain-HTTP traffic. Android blocks it by default, and Ollama doesn't use HTTPS on a home network, so without this the TV couldn't connect.
Setting up the Mac
sh
brew install ollama
ollama pull gemma4 # or gemma4:26b — fits in your 36 GB and answers better
OLLAMA_HOST=[REDACTED] ollama serve # listen on the LAN, not just localhost
Then put the Mac's IP into the TV's Settings.
For your post: the photos stay in the house, there's no API key or bill, and you can switch models by changing one setting. That covers both "keep data off a server they don't control" and "cost nothing to run". The one catch is that the Mac has to be on while you use it.
Do you want me to install Ollama and download gemma4 (about 10 GB) so I can test a real question end to end before we commit?
Sources:
- How to Run Gemma 4 with Ollama (codersera)
- Best Ollama Models 2026 (morphllm)
can you do that mac setup
$ brew install ollama 2>&1 | tail -5; which ollama; ollama --version
==> ollama
To start ollama now and restart at login:
brew services start ollama
Or, if you don't want/need a background service you can just run:
OLLAMA_FLASH_ATTENTION="1" OLLAMA_KV_CACHE_TYPE="q8_0" /opt/homebrew/opt/ollama/bin/ollama serve
/opt/homebrew/bin/ollama
Warning: could not connect to a running Ollama instance
Warning: client version is 0.35.1
$ brew services start ollama 2>&1 | tail -2; P=~/Library/LaunchAgents/homebrew.mxcl.ollama.plist; sleep 2; /usr/libexec/PlistBuddy -c "Print :EnvironmentVariables" $P
Exit code 1
==> Successfully started `ollama` (label: sh.brew.ollama)
Print: Entry, ":EnvironmentVariables", Does Not Exist
File Doesn't Exist, Will Create: ~/Library/LaunchAgents/homebrew.mxcl.ollama.plist
$ ls ~/Library/LaunchAgents/ | grep -i ollama; P=~/Library/LaunchAgents/sh.brew.ollama.plist; cat $P
sh.brew.ollama.plist
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>EnvironmentVariables</key>
<dict>
<key>OLLAMA_FLASH_ATTENTION</key>
<string>1</string>
<key>OLLAMA_KV_CACHE_TYPE</key>
<string>q8_0</string>
</dict>
<key>KeepAlive</key>
<true/>
<key>Label</key>
<string>sh.brew.ollama</string>
<key>LimitLoadToSessionType</key>
<array>
<string>Aqua</string>
<string>Background</string>
<string>LoginWindow</string>
<string>StandardIO</string>
<string>System</string>
</array>
<key>ProgramArguments</key>
<array>
<string>/opt/homebrew/opt/ollama/bin/ollama</string>
<string>serve</string>
</array>
<key>RunAtLoad</key>
<true/>
<key>StandardErrorPath</key>
<string>/opt/homebrew/var/log/ollama.log</string>
<key>StandardOutPath</key>
<string>/opt/homebrew/var/log/ollama.log</string>
<key>WorkingDirectory</key>
<string>/opt/homebrew/var</string>
</dict>
</plist>
Ollama is running as a background service. Next I'm making it listen on the home network so the TV can reach it (by default it only accepts connections from the Mac itself):
$ P=~/Library/LaunchAgents/sh.brew.ollama.plist; /usr/libexec/PlistBuddy -c "Add :EnvironmentVariables:OLLAMA_HOST string 0.0.0.0" $P && launchctl bootout gui/$(id -u)/sh.brew.ollama; sleep 1; launchctl bootstrap gui/$(id -u) $P; sleep 3; lsof -nP -iTCP:11434 -sTCP:LISTEN; curl -s localhost:11434/api/version
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NODE NAME
ollama 35387 souvikbiswas 3u IPv6 0xa12519e1d1b6e03 0t0 TCP *:11434 (LISTEN)
{"version":"0.35.1"}
Ollama is now listening on all interfaces. Downloading gemma4 (~10 GB) next:
$ ollama pull gemma4 2>&1 | tail -3; ollama list
verifying sha256 digest [K
writing manifest [K
success [K[?25h[?2026l
NAME ID SIZE MODIFIED
gemma4:latest dc35e8d9c606 6.6 GB Less than a second ago
Model downloaded. Now I'll send it the same request the TV sends, with a photo, through the Mac's network address instead of localhost:
$ IP=$(ipconfig getifaddr en0 || ipconfig getifaddr en1); echo "IP=$IP"; S=<scratch>; IMG=$(ls /System/Library/Desktop\ Pictures/*.heic 2>/dev/null | head -1); echo "$IMG"; sips -s format jpeg -Z 1024 "$IMG" --out $S/t.jpg >/dev/null; python3 - "$IP" "$S/t.jpg" <<'EOF'
import json,base64,sys,urllib.request,time
ip,img=sys.argv[1:]
b=base64.b64encode(open(img,'rb').read()).decode()
sysmsg='You help someone look through their own photos and videos on a TV. Answer about the attached picture in plain text (no markdown), concisely: usually 2–5 sentences, readable from a couch. Use the date, place and folder given as context when relevant. If you are unsure, say so briefly.'
msgs=[{'role':'system','content':sysmsg},{'role':'user','content':'Taken: 12 Mar 2024. Album/folder: Wallpapers. File: t.jpg.\n\nWhat is this?','images':[b]}]
def call(m,model='gemma4'):
t=time.time()
r=urllib.request.urlopen(urllib.request.Request(f'http://{ip}:11434/api/chat',json.dumps({'model':model,'messages':m,'stream':False,'keep_alive':'30m'}).encode(),{'Content-Type':'application/json'}))
j=json.load(r); print(f'[{time.time()-t:.1f}s]',j['message']['content']); return j['message']['content']
a=call(msgs)
msgs+= [{'role':'assistant','content':a},{'role':'user','content':'What colours stand out?'}]
call(msgs)
try: call(msgs[:2],'gemma4:nope')
except urllib.error.HTTPError as e: print(e.code, e.read().decode())
EOF
IP=192.168.68.56
/System/Library/Desktop Pictures/Mac Blue.heic
[149.4s] This is an abstract, modern digital wallpaper featuring a repeating pattern of rounded shapes. The image uses a beautiful color gradient, primarily composed of cool tones like deep blue, indigo, and lavender, which blend softly into warm peach and light pink. These smooth, stacked vertical columns and circles create a visually calming and aesthetic texture. This image, filed under Wallpapers, was captured on 12 Mar 2024.
[9.5s] The most striking colors are the deep, saturated indigo and navy blues that define the taller column shapes. These rich, cool tones stand out powerfully against the softer, lighter hues. Specifically, the gentle warm peaches and lavenders create a strong contrast with the dark blues. This beautiful blend of intense blues and delicate pastels gives the wallpaper its visually appealing, glowing quality.
404 {"error":"model 'gemma4:nope' not found"}
The test went through before your reply arrived. Since you've accepted the dialog, I'll run it again to check the first-answer time:
$ S=<scratch>; python3 - 192.168.68.56 $S/t.jpg <<'EOF'
import json,base64,sys,urllib.request,time
ip,img=sys.argv[1:]
b=base64.b64encode(open(img,'rb').read()).decode()
t=time.time()
r=urllib.request.urlopen(urllib.request.Request(f'http://{ip}:11434/api/chat',json.dumps({'model':'gemma4','messages':[{'role':'user','content':'Describe this photo in two sentences.','images':[b]}],'stream':False,'keep_alive':'30m'}).encode(),{'Content-Type':'application/json'}))
j=json.load(r); print(f'[{time.time()-t:.1f}s]',j['message']['content'])
EOF
ollama ps; /usr/libexec/ApplicationFirewall/socketfilterfw --getglobalstate; /usr/libexec/ApplicationFirewall/socketfilterfw --getappblocked /opt/homebrew/Cellar/ollama/*/bin/ollama 2>/dev/null
[10.6s] This is an abstract, seamless pattern composed of various rounded geometric shapes, such as circles and elongated rectangles. The design utilizes soft, luminous gradients of blues, purples, and peachy pastels, giving the image a dreamy and ethereal quality.
NAME ID SIZE PROCESSOR CONTEXT UNTIL
gemma4:latest dc35e8d9c606 404 MB 100% GPU 32768 29 minutes from now
Firewall is enabled. (State = 1)
Incoming connection to /opt/homebrew/Cellar/ollama/0.35.1_1/bin/ollama is permitted.
Your Mac is set up, and Gemma answers questions about a photo over the home network.
What I did
- Installed Ollama 0.35.1 with Homebrew. It runs as a background service that starts when you log in.
- Set OLLAMA_HOST=[REDACTED] in the service's launch file (~/Library/LaunchAgents/sh.brew.ollama.plist), so it accepts connections from other devices on the Wi-Fi.
- Downloaded gemma4 (6.6 GB) and checked that it runs on the M3 Max's GPU.
- The dialog you accepted was the macOS firewall. It now shows incoming connections to Ollama as permitted.
Test results (I sent the app's exact request to http://[REDACTED]:11434 with a 1024px JPEG):
- First question: took 149 s. That was the model loading for the first time, and probably the firewall dialog waiting on you too.
- Same question again, model already loaded: 10.6 s.
- Follow-up question: 9.5 s, and the answer used the earlier conversation.
- Answers: accurate, plain text, 2–5 sentences as the app asks for, and they used the folder and date I passed in.
- Model name that doesn't exist: returns 404 with model '…' not found, which the app turns into the "run ollama pull …" message.
These requests came from the Mac to its own network address. The TV itself hasn't connected yet.
On the TV: go to Settings → Gemma server and enter [REDACTED]. The app fills in the port.
Two things to watch
- The Mac's IP can change. If the router hands out a new address, the TV can't find Gemma. Set a fixed (reserved) address for the Mac in your router, or use its .local name.
- brew services restart ollama undoes the network setting. It rewrites the launch file without OLLAMA_HOST, so after it Ollama only listens to the Mac itself. Plain reboots are fine. If you run that command, tell me and I'll put the setting back.
The Gemma code changes are still not committed. Should I commit them now, or first build the app onto the TV and test it there?
Can we use faster model to make it reply under 5 seconds or even less
I'll first measure where the 10 s goes (reading the image, thinking, or writing the answer), then compare models and settings.
$ ollama pull gemma4:e2b 2>&1 | tail -1; S=<scratch>; sips -Z 512 $S/t.jpg --out $S/t512.jpg >/dev/null; python3 - $S <<'EOF'
import json,base64,sys,urllib.request,time
S=sys.argv[1]
sysmsg='You help someone look through their own photos and videos on a TV. Answer about the attached picture in plain text (no markdown), concisely: usually 2–5 sentences, readable from a couch. Use the date, place and folder given as context when relevant. If you are unsure, say so briefly.'
def call(model,img,think=None,label=''):
b=base64.b64encode(open(f'{S}/{img}','rb').read()).decode()
body={'model':model,'messages':[{'role':'system','content':sysmsg},{'role':'user','content':'Taken: 12 Mar 2024. Album/folder: Wallpapers.\n\nWhat is this?','images':[b]}],'stream':False,'keep_alive':'30m'}
if think is not None: body['think']=think
t=time.time()
try:
j=json.load(urllib.request.urlopen(urllib.request.Request('http://localhost:11434/api/chat',json.dumps(body).encode(),{'Content-Type':'application/json'})))
except urllib.error.HTTPError as e: print(label,e.code,e.read().decode()); return
ns=1e9
print(f"{label:28} total {time.time()-t:5.1f}s | load {j.get('load_duration',0)/ns:4.1f} prompt {j.get('prompt_eval_count')}tok {j.get('prompt_eval_duration',0)/ns:4.1f}s | gen {j.get('eval_count')}tok {j.get('eval_duration',0)/ns:4.1f}s | thinking {len(j['message'].get('thinking') or '')}ch")
for m in ['gemma4','gemma4:e2b']:
call(m,'t.jpg',label=f'{m} warmup')
call(m,'t.jpg',label=f'{m} 1024 default')
call(m,'t.jpg',think=False,label=f'{m} 1024 think=false')
call(m,'t512.jpg',think=False,label=f'{m} 512 think=false')
EOF
success [K[?25h[?2026l
gemma4 warmup total 27.7s | load 0.0 prompt 545tok 2.2s | gen 521tok 25.4s | thinking 1654ch
gemma4 1024 default total 29.2s | load 0.0 prompt 545tok 0.1s | gen 563tok 29.0s | thinking 1797ch
gemma4 1024 think=false total 5.2s | load 0.0 prompt 543tok 2.4s | gen 46tok 2.8s | thinking 0ch
gemma4 512 think=false total 4.9s | load 0.0 prompt 223tok 1.0s | gen 61tok 3.9s | thinking 0ch
gemma4:e2b warmup total 15.3s | load 5.1 prompt 545tok 1.5s | gen 369tok 8.7s | thinking 1151ch
gemma4:e2b 1024 default total 9.9s | load 0.0 prompt 545tok 0.0s | gen 400tok 9.8s | thinking 1401ch
gemma4:e2b 1024 think=false total 3.2s | load 0.0 prompt 543tok 1.6s | gen 56tok 1.5s | thinking 0ch
gemma4:e2b 512 think=false total 2.7s | load 0.0 prompt 223tok 0.5s | gen 71tok 2.1s | thinking 0ch
$ S=<scratch>; F=$(find ~/Pictures ~/Downloads ~/Desktop -maxdepth 3 -iname '*.jp*g' -size +300k 2>/dev/null | head -1); echo "$F"; sips -s format jpeg -Z 1024 "$F" --out $S/real.jpg >/dev/null && python3 - $S <<'EOF'
import json,base64,sys,urllib.request,time
S=sys.argv[1]
sysmsg='You help someone look through their own photos and videos on a TV. Answer about the attached picture in plain text (no markdown), concisely: usually 2–5 sentences, readable from a couch. Use the date, place and folder given as context when relevant. If you are unsure, say so briefly.'
b=base64.b64encode(open(f'{S}/real.jpg','rb').read()).decode()
for m in ['gemma4','gemma4:e2b']:
for q in ['What is this?','Where was this taken?']:
t=time.time()
j=json.load(urllib.request.urlopen(urllib.request.Request('http://localhost:11434/api/chat',json.dumps({'model':m,'think':False,'messages':[{'role':'system','content':sysmsg},{'role':'user','content':q,'images':[b]}],'stream':False,'keep_alive':'30m'}).encode(),{'Content-Type':'application/json'})))
print(f'--- {m} | {q} [{time.time()-t:.1f}s]\n{j["message"]["content"]}')
EOF
~/Pictures/story_4.jpg
--- gemma4 | What is this? [10.9s]
This is a beautiful shot of a large ocean wave crashing. The water has a vibrant turquoise color where the wave is breaking against the deep dark blue of the open sea. It looks like a dynamic scene from the beach.
--- gemma4 | Where was this taken? [4.7s]
This appears to be a striking aerial or high-angle shot of a wave rolling into the ocean. The image shows bright white foam contrasting sharply with the deep turquoise and dark blue of the water. Without specific metadata, I cannot tell you where this photo was taken.
--- gemma4:e2b | What is this? [6.5s]
This appears to be a video or photo of ocean waves crashing onto a shore. The water is a deep turquoise color, contrasting with the white foam of the breaking waves. It looks like a dynamic scene capturing the energy of the ocean.
--- gemma4:e2b | Where was this taken? [2.5s]
This image appears to be a long exposure of ocean waves crashing onto a shore. It shows white, foamy waves meeting dark blue water. Without specific date, place, or folder information, it is impossible to say exactly where this was taken.
$ grep -n "Future<String> ask\|final body = jsonEncode\|'stream': false\|final (status, json)\|/// One /api/chat call\|/// The photo (or video poster)" lib/services/gemma_service.dart
59: Future<String> ask(String question) async {
87: final body = jsonEncode({
101: 'stream': false,
107: final (status, json) = await _post(server, body);
127: /// One /api/chat call: (HTTP status, decoded JSON body).
154: /// The photo (or video poster) as a ~1024px JPEG, base64-encoded.
$ sed -n 55,60p lib/services/gemma_service.dart; sed -n 96,153p lib/services/gemma_service.dart
static bool get hasServer => SettingsController.instance.gemmaServer != null;
/// Asks [question] about the photo; returns the answer text. Throws a
/// [GemmaException] with a user-facing message on failure.
Future<String> ask(String question) async {
final settings = SettingsController.instance;
'when relevant. If you are unsure, say so briefly.',
},
..._history,
turn,
],
'stream': false,
// Keep the model loaded between questions while browsing (loading it
// takes several seconds; Ollama's default unloads after 5 minutes).
'keep_alive': '30m',
});
final (status, json) = await _post(server, body);
final error = json['error'] as String?;
if (status == 200 && error == null) {
final answer =
((json['message'] as Map?)?['content'] as String? ?? '').trim();
if (answer.isEmpty) {
throw const GemmaException('Gemma didn’t return an answer.');
}
_history
..add(turn)
..add({'role': 'assistant', 'content': answer});
return answer;
}
throw GemmaException(switch (status) {
404 => 'The model “$model” isn’t on the Gemma server yet — run '
'“ollama pull $model” there, or pick another model in Settings.',
_ => 'Gemma error $status${error == null ? '' : ': $error'}',
});
}
/// One /api/chat call: (HTTP status, decoded JSON body).
Future<(int, Map<String, dynamic>)> _post(String server, String
…[1096 more chars]
$ python3 - <<'EOF'
p='lib/services/gemma_service.dart'
s=open(p).read()
def r(a,b):
global s
assert a in s, a
s=s.replace(a,b)
r(''' /// Asks [question] about the photo; returns the answer text. Throws a
/// [GemmaException] with a user-facing message on failure.
Future<String> ask(String question) async {''',''' /// Asks [question] about the photo; returns the answer text. The answer
/// streams in: [onPartial] gets the text so far as each piece arrives, so
/// it can be shown while Gemma is still writing. Throws a [GemmaException]
/// with a user-facing message on failure.
Future<String> ask(
String question, {
void Function(String partial)? onPartial,
}) async {''')
a=s.index(" 'stream': false,")
b=s.index(" /// The photo (or video poster) as a ~1024px JPEG")
s=s[:a]+''' 'stream': true,
// Answer straight away: with thinking on, Gemma 4 reasons for ~500
// tokens first (10–30 s), and describing a photo doesn't need it.
'think': false,
// Keep the model loaded between questions while browsing (loading it
// takes several seconds; Ollama's default unloads after 5 minutes).
'keep_alive': '30m',
});
final answer = (await _chat(server, model, body, onPartial)).trim();
if (answer.isEmpty) {
throw const GemmaException('Gemma didn’t return an answer.');
}
_history
..add(turn)
..add({'role': 'assistant', 'content': answer});
return answer;
}
/// One streaming /api/chat call; returns the whole answer. Ollama sends a
/// JSON object per line, each with the next piece of the message.
Future<String> _chat(
String server,
String model,
String body,
void Function(String partial)? onPartial,
) async {
final client = HttpClient()..connectionTimeout = const Duration(seconds: 8);
const unreachable = GemmaException(
'Couldn’t reach the Gemma server — is the computer on, with Ollama '
'running and listening on the network (OLLAMA_HOST=0.0.0.0)?');
try {
final req = await client.postUrl(Uri.parse('$server/api/chat'));
req.headers.contentType = ContentType.json;
req.add(utf8.encode(body));
// Generous: the first question may wait for the model to load.
final res = await req.close().timeout(const Duration(seconds: 180));
final lines = res
.transform(utf8.decoder)
.transform(const LineSplitter())
.where((l) => l.trim().isNotEmpty)
.timeout(const Duration(seconds: 60));
final answer = StringBuffer();
await for (final line in lines) {
final json = jsonDecode(line) as Map<String, dynamic>;
final error = json['error'] as String?;
if (res.statusCode != 200 || error != null) {
throw GemmaException(switch (res.statusCode) {
404 => 'The model “$model” isn’t on the Gemma server yet — run '
'“ollama pull $model” there, or pick another model in '
…[1798 more chars]
2:import 'dart:convert';
145: .transform(const LineSplitter())
lib/screens/settings_screen.dart:251: hint: 'Ollama tag, e.g. gemma4 or gemma4:26b',
$ python3 - <<'EOF'
p='lib/widgets/ask_panel.dart'
s=open(p).read()
old=''' } on GemmaException catch (e) {
turn.error = e.message;
} catch (e) {
turn.error = 'Something went wrong: $e';
}'''
assert old in s
s=s.replace(old,''' } on GemmaException catch (e) {
turn
..answer = null // drop a half-streamed answer
..error = e.message;
} catch (e) {
turn
..answer = null
..error = 'Something went wrong: $e';
}''')
open(p,'w').write(s)
p='lib/screens/settings_screen.dart'
s=open(p).read()
s=s.replace("hint: 'Ollama tag, e.g. gemma4 or gemma4:26b',","hint: 'Ollama tag — gemma4:e2b (fastest) or gemma4',")
open(p,'w').write(s)
EOF
flutter analyze 2>&1 | tail -3; git checkout analysis_options.yaml 2>/dev/null; git status --short
Upgrading analysis_options.yaml to exclude build and platform directories.
Analyzing tv_gallery_gemma...
No issues found! (ran in 7.9s)
M android/app/src/main/AndroidManifest.xml
M lib/screens/settings_screen.dart
M lib/screens/slideshow_screen.dart
M lib/screens/viewer_screen.dart
M lib/services/answer_voice.dart
RM lib/services/gemini_service.dart -> lib/services/gemma_service.dart
M lib/services/settings_controller.dart
M lib/widgets/ask_panel.dart
M lib/widgets/zoomable_image.dart
Code updated and analyzer clean. Now I'm timing the streaming request against the real Gemma server: when the first words arrive and when the answer finishes.
$ S=<scratch>; cat > $S/stream_test.dart <<'EOF'
import 'dart:convert';
import 'dart:io';
Future<void> main(List<String> a) async {
final img = base64Encode(File(a[1]).readAsBytesSync());
for (final model in [a[2], a[2]]) {
final sw = Stopwatch()..start();
final client = HttpClient();
final req = await client.postUrl(Uri.parse('${a[0]}/api/chat'));
req.headers.contentType = ContentType.json;
req.add(utf8.encode(jsonEncode({
'model': model, 'stream': true, 'think': false, 'keep_alive': '30m',
'messages': [{'role': 'user', 'content': 'What is this?', 'images': [img]}],
})));
final res = await req.close();
final answer = StringBuffer();
int? first;
await for (final line in res.transform(utf8.decoder).transform(const LineSplitter()).where((l) => l.trim().isNotEmpty)) {
final j = jsonDecode(line) as Map<String, dynamic>;
if (j['error'] != null) { print('${res.statusCode} ${j['error']}'); break; }
final piece = (j['message'] as Map?)?['content'] as String?;
if (piece != null && piece.isNotEmpty) { first ??= sw.elapsedMilliseconds; answer.write(piece); }
if (j['done'] == true) break;
}
client.close();
print('$model first words ${first}ms, done ${sw.elapsedMilliseconds}ms: ${answer.toString().trim()}');
}
}
EOF
dart run $S/stream_test.dart http://192.168.68.56:11434 $S/real.jpg gemma4:e2b; dart run $S/stream_test.dart http://192.168.68.56:11434 $S/real.jpg gemma4; dart run $S/stream_test.dart http://192.168.68.56:11434 $S/real.jpg nope 2>&1 | head -1
gemma4:e2b first words 847ms, done 6094ms: This is a photograph of a **seascape** or **ocean wave breaking on a shore**.
Here's a breakdown of what the image shows:
* **Waves:** The foreground is dominated by white, foamy waves crashing onto the shore.
* **Water Color:** The water in the deeper areas appears very dark (almost black), suggesting deep water or a dark sky/lighting condition. However, where the water is shallower or illuminated by light, it displays beautiful shades of **turquoise and deep blue/teal**.
* **Motion/Texture:** The photograph captures the dynamic motion of the water—the energy of the breaking waves and the flow of the ocean.
* **Contrast:** There is a very high contrast between the bright white of the surf and the dark, deep tones of the water.
In short, it's a dramatic and beautiful shot of the **ocean waves**.
gemma4:e2b first words 89ms, done 4384ms: This is a **photograph of ocean waves crashing onto a shore**.
Here's a breakdown of what the image shows:
* **Waves:** There is a dynamic interaction between the ocean water and the land. You can see white, foamy waves breaking and rolling in, contrasted against the deep blue/teal color of the water.
* **Water Color:** The water exhibits beautiful shades of turquoise, teal, and deep navy/black in the deeper sections, suggesting clear, deep ocean water.
* **Action:** The photograph captures the energy of the surf, showing the white water spray and the continuous movement of the waves.
**In summa
…[1613 more chars]
Answers are now about 2–3 s from asking to the full answer on screen with the default gemma4:e2b, and the first words show up in under a second. All of this was measured on the Mac against its own network address. I haven't run it on the TV yet.
What I changed
- Thinking off: this was the main cause. Gemma 4 was silently reasoning for about 500 words before every answer, which took 10–29 s. Turning it off (think: false) costs almost nothing in quality for describing a photo.
- Default model is now gemma4:e2b: I downloaded it next to gemma4, which stays one setting away under Settings → Model if you prefer slightly better wording.
- Answers stream in: the words appear on screen as Gemma writes them, instead of all at once at the end. If something fails halfway through, the partial answer is replaced by the error message. Reading aloud still starts once the full answer is in.
Measured (photo already prepared, model loaded)
| Full answer | First words | |
|---|---|---|
gemma4:e2b |
~2.5–3 s | ~0.1–0.9 s |
gemma4 |
~5 s | ~0.1–0.9 s* |
| Before (thinking on) | 10–29 s | same as full answer |
*The first gemma4 question after using e2b took about 6 s before any words appeared, because the Mac had to load the model. That only happens when you switch models or after 30 minutes idle.
In the streaming test I left out the app's short-answer instruction, so the answers ran long and slow there (4–13 s). With the instruction, as in the earlier test, they're 2–5 sentences, which gives the times in the table.
flutter analyze passes. Should I commit now, or build onto the TV first so you can try it there?
build and install on the TV
$ adb devices -l; cat ~/MobileDev/Flutter/tv_gallery_gemma/.fvmrc; ls ~/.claude/projects/-Users-souvikbiswas-MobileDev-Flutter-tv-gallery-gemma/memory/ 2>/dev/null
List of devices attached
{
"flutter": "stable"
}
let's do debug for now
No TV is connected over adb right now. I'll check whether it shows up on the network:
$ adb mdns check; adb mdns services; grep -rh "adb connect" ~/.zsh_history 2>/dev/null | tail -3
mdns daemon version [adb discovery 0.0.0]
List of discovered mdns services
$ arp -a | grep -v incomplete | head -30; for ip in $(arp -a | grep -v incomplete | sed -E 's/.*\(([0-9.]+)\).*/\1/'); do (nc -z -G 1 $ip 5555 2>/dev/null && echo "adb port open: $ip") & done; wait
? (192.168.68.1) at 48:22:54:c3:8a:e8 on en0 ifscope [ethernet]
? (192.168.68.50) at 7c:9e:bd:6:b0:74 on en0 ifscope [ethernet]
? (192.168.68.51) at 90:70:69:12:a6:a4 on en0 ifscope [ethernet]
? (192.168.68.52) at 10:5a:95:3f:f4:36 on en0 ifscope [ethernet]
? (192.168.68.55) at da:6e:e6:4a:5e:71 on en0 ifscope [ethernet]
? (192.168.68.56) at 60:3e:5f:88:47:44 on en0 ifscope permanent [ethernet]
? (192.168.68.58) at ba:7:f0:f7:3f:c1 on en0 ifscope [ethernet]
? (192.168.68.59) at 38:e1:3d:ba:53:63 on en0 ifscope [ethernet]
? (192.168.68.60) at 32:8a:d5:d:97:ae on en0 ifscope [ethernet]
? (192.168.68.62) at 78:46:5c:60:b0:d3 on en0 ifscope [ethernet]
? (192.168.68.63) at 84:d3:28:da:69:4d on en0 ifscope [ethernet]
? (192.168.68.65) at 68:ef:dc:9a:e6:e2 on en0 ifscope [ethernet]
? (192.168.68.66) at a4:83:e7:e4:e1:89 on en0 ifscope [ethernet]
? (192.168.68.67) at 56:c4:3d:e6:bb:3b on en0 ifscope [ethernet]
? (192.168.71.255) at ff:ff:ff:ff:ff:ff on en0 ifscope [ethernet]
mdns.mcast.net (224.0.0.251) at 1:0:5e:0:0:fb on en0 ifscope permanent [ethernet]
? (239.255.255.250) at 1:0:5e:7f:ff:fa on en0 ifscope permanent [ethernet]
I can't find the TV on the network. None of the devices answers on adb's port (5555), and nothing advertises wireless debugging. I need you to turn on debugging on the Bravia:
- Go to Settings → System → About and click Build seven times to unlock developer options (skip this if you've done it before).
- Go to Settings → System → Developer options and turn on USB debugging, plus Network debugging / ADB over network if it's there.
- Send me the TV's IP address from Settings → Network & Internet → (your Wi-Fi) → IP address.
Once you send the IP, I'll run adb connect <ip>. You'll need to accept an "Allow USB debugging?" prompt on the TV, then I'll start the debug build with flutter run. If you'd rather connect it yourself, type ! adb connect <tv-ip> here and I'll take it from there.
chedck now
$ adb devices -l; adb mdns services; for ip in $(arp -a | grep -v incomplete | sed -E 's/.*\(([0-9.]+)\).*/\1/'); do (nc -z -G 1 $ip 5555 2>/dev/null && echo "adb port open: $ip") & done; wait
List of devices attached
<tv-adb-serial> device product:BRAVIA_AE2_PA model:BRAVIA_4K_AE2 device:BRAVIA_AE2 transport_id:759
List of discovered mdns services
<tv-adb-serial> _adb-tls-connect._tcp 192.168.68.62:39273
The TV is connected (BRAVIA 4K AE2 at [REDACTED]). Starting the debug build on it:
$ S=<scratch> && mkfifo $S/run.fifo 2>/dev/null; (tail -f $S/run.fifo | flutter run --debug -d "<tv-adb-serial>" > $S/run.log 2>&1) &
echo started
Command running in background with ID: bdfva43ve. Output is being written to: <tmp> You will be notified when it completes. To check interim output, use Read on that file path.
{
"query": "select:Monitor",
"max_results": 1
}
[image]
$ S=<scratch>; until grep -qE "Flutter run key commands|Error|FAILURE|Exception|No supported devices|not found" $S/run.log; do sleep 2; done; tail -25 $S/run.log
I/flutter ( 7444): These options are going to go away in an upcoming Flutter release. Remove
I/flutter ( 7444): the explicit opt-out. If you need to opt-out, please report a bug describing
I/flutter ( 7444): the issue.
I/flutter ( 7444):
I/flutter ( 7444): https://github.com/flutter/flutter/issues/new?template=02_bug.yml
I/flutter ( 7444):
D/FlutterRenderer( 7444): Width is zero. 0,0
Syncing files to device BRAVIA 4K AE2 (wireless)...
I/Choreographer( 7444): Skipped 283 frames! The application may be doing too much work on its main thread.
Syncing files to device BRAVIA 4K AE2 (wireless)... 168ms
Flutter run key commands.
r Hot reload. 🔥🔥🔥
R Hot restart.
h List all available interactive commands.
d Detach (terminate "flutter run" but leave application running).
c Clear the screen
q Quit (terminate the application on the device).
A Dart VM Service on BRAVIA 4K AE2 is available at: http://127.0.0.1:50620/lBlMT0wSeBI=/
The Flutter DevTools debugger and profiler on BRAVIA 4K AE2 is available at: http://127.0.0.1:50620/lBlMT0wSeBI=/devtools/?uri=ws://127.0.0.1:50620/lBlMT0wSeBI=/ws
D/FlutterRenderer( 7444): Width is zero. 0,0
D/FlutterJNI( 7444): Sending viewport metrics to the engine.
I/TextToSpeech( 7444): Setting up the connection to TTS engine...
E/OpenGLRenderer( 7444): Unable to match the desired swap behavior.
tail -n 0 -f <scratch>/run.log | grep -E --line-buffered "EXCEPTION|Exception|Error:|══╡|Lost connection|Application finished"
Monitor started (task b1r1vuvr2, expires in 30m unless the source ends first; you get one notice at expiry — re-arm if you still need the watch). You will be notified on each event. Keep working — do not poll or sleep. Events may arrive while you are waiting for the user — an event is not their reply.
The debug build is running on the TV, the BRAVIA 4K AE2, connected over Wi-Fi. I'm watching its log for errors while you try it.
To try it:
1. In Memories, open Settings → Gemma, go to Gemma server → Add, and enter [REDACTED] (the Mac). Leave Model as gemma4:e2b.
2. Open a photo and press Down on the remote to open Ask Gemma, then pick a suggestion or ask something.
The first question can take about 5 s while the Mac loads the model. After that, answers should start appearing within about a second.
Tell me how it goes. I can hot-reload any changes straight onto the TV.
'~/Desktop/Screenshots/tv-125724.png'
remove the "-- or use the keyboard mic" text part in hint
~/Desktop/Screenshots/tv-125724.png
[image]
$ sed -i '' "s/'Ask anything about this photo — or use the keyboard’s mic'/'Ask anything about this photo'/" lib/widgets/ask_panel.dart && grep -n "Ask anything" lib/widgets/ask_panel.dart && echo r > <scratch>/run.fifo; sleep 4; tail -3 <scratch>/run.log
353: ? 'Ask anything about this photo'
Performing hot reload...
Reloaded 1 of 2219 libraries in 3,294ms (compile: 80 ms, reload: 1318 ms, reassemble: 1648 ms).
The hint now just says "Ask anything about this photo". I hot-reloaded it onto the TV, so it shows next time Ask Gemma opens.
seems like even with the keyboard closed, I can't hit down to selct any of the options from there
Down is being swallowed by the text field: it handles Down itself, as "move the cursor to the end". Once there's a conversation, Down already moves into the answers. Before any question, it should go to the suggestion pills instead:
$ python3 - <<'EOF'
p='lib/widgets/ask_panel.dart'
s=open(p).read()
old=''' // Down (keyboard hidden) moves into the answers to scroll them.
if (event.logicalKey == LogicalKeyboardKey.arrowDown &&
_field.hasFocus &&
_turns.isNotEmpty) {
if (event is KeyDownEvent) _answers.requestFocus();
return KeyEventResult.handled;
}'''
assert old in s
s=s.replace(old,''' // Down (keyboard hidden) moves into the answers to scroll them, or
// to the suggestions before anything's been asked. (The field would
// otherwise keep Down to move its cursor.)
if (event.logicalKey == LogicalKeyboardKey.arrowDown &&
_field.hasFocus) {
if (event is KeyDownEvent) {
if (_turns.isNotEmpty) {
_answers.requestFocus();
} else {
_field.focusInDirection(TraversalDirection.down);
}
}
return KeyEventResult.handled;
}''')
open(p,'w').write(s)
EOF
echo r > <scratch>/run.fifo; sleep 4; tail -1 <scratch>/run.log
Reloaded 1 of 2219 libraries in 2,909ms (compile: 19 ms, reload: 1150 ms, reassemble: 1587 ms).
I fixed it and hot-reloaded it onto the TV. With the keyboard closed and nothing asked yet, Down from the question field now jumps to the suggestion pills, and Up should take you back to the field. After you've asked something, Down still moves into the answers to scroll them. Close and reopen Ask Gemma to try it.
closing the keyboard with back shoudl actually unfocus the text field as well I think
First I'll check how the TV reports the keyboard closing. I can drive the TV over adb and take screenshots:
$ cat android/app/src/main/kotlin/*/*/*/MainActivity.kt 2>/dev/null | head -60 || find android/app/src/main -name "*.kt" | head; adb shell dumpsys input_method | grep -E "mInputShown|mImeWindowVis|mCurMethodId|mShowRequested" | head
package com.souvikbiswas.tvgallery
import io.flutter.embedding.android.FlutterActivity
import io.flutter.embedding.engine.FlutterEngine
class MainActivity : FlutterActivity() {
override fun configureFlutterEngine(flutterEngine: FlutterEngine) {
super.configureFlutterEngine(flutterEngine)
flutterEngine.platformViewsController.registry.registerViewFactory(
"tv_gallery/video",
VideoPlayerViewFactory(flutterEngine.dartExecutor.binaryMessenger),
)
MediaIndexChannel(applicationContext, flutterEngine.dartExecutor.binaryMessenger)
}
}
mCurMethodId=com.google.android.inputmethod.latin/com.android.inputmethod.latin.LatinIME
mInputShown=true
$ grep -n "class _AskPanelState\|void initState() {\| void dispose() {" lib/widgets/ask_panel.dart | head -4
58:class _AskPanelState extends State<_AskPanel> {
71: void initState() {
97: void dispose() {
504: void initState() {
$ python3 - <<'EOF'
p='lib/widgets/ask_panel.dart'
s=open(p).read()
s=s.replace('''class _AskPanelState extends State<_AskPanel> {''','''class _AskPanelState extends State<_AskPanel> with WidgetsBindingObserver {
@override
void didChangeMetrics() {
final v = View.of(context);
debugPrint('DBGIME insets=${v.viewInsets.bottom} padding=${v.viewPadding.bottom} size=${v.physicalSize}');
}
''',1)
s=s.replace(''' void initState() {
super.initState();
_loadBackdrop();''',''' void initState() {
super.initState();
WidgetsBinding.instance.addObserver(this);
_loadBackdrop();''',1)
open(p,'w').write(s)
EOF
S=<scratch>; echo R > $S/run.fifo; sleep 6; tail -2 $S/run.log; adb exec-out screencap -p > $S/s1.png
Performing hot restart...
$ S=<scratch>; sleep 3; tail -2 $S/run.log; adb exec-out screencap -p > $S/s1.png; sips -Z 960 $S/s1.png >/dev/null
D/FlutterJNI( 7444): Sending viewport metrics to the engine.
I/iswas.tvgallery( 7444): Background concurrent copying GC freed 59728(1758KB) AllocSpace objects, 0(0B) LOS objects, 49% free, 3417KB/6834KB, paused 127us,42us total 131.333ms
<scratch>/s1.png
[image]
$ S=<scratch>; sleep 4; adb exec-out screencap -p > $S/s1.png; sips -Z 960 $S/s1.png >/dev/null
(Bash completed with no output)
<scratch>/s1.png
[image]
$ S=<scratch>; adb shell input keyevent DPAD_DOWN; sleep 3; adb shell dumpsys input_method | grep mInputShown; grep DBGIME $S/run.log | tail -3; echo ---back; adb shell input keyevent BACK; sleep 2; adb shell dumpsys input_method | grep mInputShown; grep DBGIME $S/run.log | tail -3; adb exec-out screencap -p > $S/s2.png; sips -Z 960 $S/s2.png >/dev/null
mInputShown=false
I/flutter ( 7444): DBGIME insets=0.0 padding=0.0 size=Size(1920.0, 1080.0)
I/flutter ( 7444): DBGIME insets=0.0 padding=0.0 size=Size(1920.0, 1080.0)
I/flutter ( 7444): DBGIME insets=0.0 padding=0.0 size=Size(1920.0, 1080.0)
---back
mInputShown=false
I/flutter ( 7444): DBGIME insets=0.0 padding=0.0 size=Size(1920.0, 1080.0)
I/flutter ( 7444): DBGIME insets=0.0 padding=0.0 size=Size(1920.0, 1080.0)
I/flutter ( 7444): DBGIME insets=0.0 padding=0.0 size=Size(1920.0, 1080.0)
<scratch>/s2.png
[image]
I'm checking how the TV signals the keyboard closing. A debug probe I added caused an error (I didn't unregister it when the panel closes), so I'm looking at that first:
$ S=<scratch>; grep -n -A12 "EXCEPTION CAUGHT" $S/run.log | tail -14
432:══╡ EXCEPTION CAUGHT BY WIDGETS LIBRARY ╞═══════════════════════════════════════════════════════════
433-The following assertion was thrown while dispatching notifications for
434-WidgetsBindingObserver.didChangeMetrics:
435-This widget has been unmounted, so the State no longer has a context (and should be considered
436-defunct).
437-Consider canceling any active work during "dispose" or using the "mounted" getter to determine if
438-the State is still active.
439-
440-When the exception was thrown, this was the stack:
441-#0 State.context.<anonymous closure> (package:flutter/src/widgets/framework.dart:952:9)
442-#1 State.context (package:flutter/src/widgets/framework.dart:958:6)
443-#2 _AskPanelState.didChangeMetrics (package:tv_gallery/widgets/ask_panel.dart:61:23)
444-#3 WidgetsBinding.handleMetricsChanged (package:flutter/src/widgets/binding.dart:926:18)
$ adb shell input keyevent DPAD_CENTER; sleep 3; adb shell dumpsys input_method | grep mInputShown; adb shell dumpsys window | grep -iE "InsetsSource.*ime|type=ime|mImeInsetsSourceProvider|isImeShowing|ImeInsetsSourceProvider" | head -8
mInputShown=true
mIsImeShowing=false
InsetsSource id=3 type=ime frame=[0,1080][1920,1080] visibleFrame=[0,1080][1920,1080] visible=true flags= insetsRoundedCornerFrame=false
InsetsSourceControl: {3 mType=ime mSurfacePosition=Point(0, 0) mInsetsHint=Insets{left=0, top=0, right=0, bottom=0}}
ImeInsetsSourceProvider
mSource=InsetsSource id=3 type=ime frame=[0,1080][1920,1080] visibleFrame=[0,1080][1920,1080] visible=true flags= insetsRoundedCornerFrame=false
mControl=InsetsSourceControl mId=3 mType=ime mLeash=Surface(name=Surface(name=73cc467 InputMethod)/@0xb81a8b2 - animation-leash of insets_animation)/@0xda3be45 mInitiallyVisible=false mSurfacePosition=Point(0, 0) mInsetsHint=Insets{left=0, top=0, right=0, bottom=0} mSkipAnimationOnce=false
$ adb shell getprop ro.build.version.sdk; cd ~/MobileDev/Flutter/tv_gallery_gemma && ls android/app/src/main/kotlin/com/souvikbiswas/tvgallery/; grep -n "dependencies" -A8 android/app/build.gradle.kts; grep -n "minSdk" -r android/app/build.gradle.kts; sed -n 1,40p android/app/src/main/kotlin/com/souvikbiswas/tvgallery/MediaIndexChannel.kt
34
MainActivity.kt
MediaIndexChannel.kt
VideoPlayerView.kt
VideoPlayerViewFactory.kt
43:dependencies {
44- // ExoPlayer (Media3) for native video rendered to a SurfaceView — gives true
45- // HDR10/HLG/Dolby Vision passthrough and smooth high-frame-rate playback,
46- // which the Flutter-texture path can't do.
47- implementation("androidx.media3:media3-exoplayer:1.4.1")
48-}
49-
50-flutter {
51- source = "../.."
android/app/build.gradle.kts:22: minSdk = flutter.minSdkVersion
package com.souvikbiswas.tvgallery
import android.content.Context
import android.database.Cursor
import android.graphics.Bitmap
import android.graphics.BitmapFactory
import android.media.ExifInterface
import android.media.ThumbnailUtils
import android.os.Build
import android.os.CancellationSignal
import android.os.Environment
import android.os.Handler
import android.os.Looper
import android.os.storage.StorageManager
import android.provider.MediaStore
import android.provider.MediaStore.Files.FileColumns
import android.util.Size
import io.flutter.plugin.common.BinaryMessenger
import io.flutter.plugin.common.MethodCall
import io.flutter.plugin.common.MethodChannel
import java.io.File
import java.io.FileOutputStream
import java.util.concurrent.ConcurrentHashMap
import java.util.concurrent.atomic.AtomicBoolean
import java.util.concurrent.Executors
/**
* Lightweight MediaStore queries for the folder browser.
*
* photo_manager's asset listing materialises a full entity per file — inclu
…[693 more chars]
Okay, I'll leave it. Removing my debug probe:
$ python3 - <<'EOF'
p='lib/widgets/ask_panel.dart'
s=open(p).read()
probe='''class _AskPanelState extends State<_AskPanel> with WidgetsBindingObserver {
@override
void didChangeMetrics() {
final v = View.of(context);
debugPrint('DBGIME insets=${v.viewInsets.bottom} padding=${v.viewPadding.bottom} size=${v.physicalSize}');
}
'''
assert probe in s
s=s.replace(probe,'class _AskPanelState extends State<_AskPanel> {\n')
s=s.replace(''' WidgetsBinding.instance.addObserver(this);
_loadBackdrop();''',''' _loadBackdrop();''')
open(p,'w').write(s)
EOF
grep -c "DBGIME\|addObserver" lib/widgets/ask_panel.dart; S=<scratch>; echo r > $S/run.fifo; sleep 4; tail -1 $S/run.log; git diff --stat
0
Reloaded 8 of 2219 libraries in 3,098ms (compile: 46 ms, reload: 1362 ms, reassemble: 1610 ms).
analysis_options.yaml | 5 +
android/app/src/main/AndroidManifest.xml | 9 +-
lib/screens/settings_screen.dart | 58 +++++---
lib/screens/slideshow_screen.dart | 4 +-
lib/screens/viewer_screen.dart | 6 +-
lib/services/answer_voice.dart | 2 +-
lib/services/gemma_service.dart | 227 ++++++++++++++++---------------
lib/services/settings_controller.dart | 55 ++++++--
lib/widgets/ask_panel.dart | 74 ++++++----
lib/widgets/zoomable_image.dart | 2 +-
10 files changed, 260 insertions(+), 182 deletions(-)
$ git checkout analysis_options.yaml && git status --short | wc -l
Updated 1 path from the index
9
I'm leaving the Back/keyboard behaviour as it is. I removed the temporary debug code (it was what caused the error in the log) and hot-reloaded the TV. I also reverted the analysis_options.yaml change that flutter run added on its own.
If you change your mind later, here's what I found: the TV keyboard never tells the app it's open (the space it covers always reads 0). Android does still mark it as visible, so the app could ask Android directly through a small piece of Android code.
Nothing is committed yet. The changes are the switch to Gemma, the speed changes, the shorter hint, and Down reaching the suggestion pills. The debug app is still running on the TV. Say when you want me to commit.
The old panels are still registered for my removed debug hook in the running app, so they keep throwing. A hot restart clears them:
$ S=<scratch>; echo R > $S/run.fifo; sleep 8; grep -n "Restarted application" $S/run.log | tail -1; L=$(grep -n "Restarted application" $S/run.log | tail -1 | cut -d: -f1); tail -n +$L $S/run.log | grep -c didChangeMetrics
710:Restarted application in 7,528ms.
0
Those alerts are from before the restart arriving late. To confirm:
$ S=<scratch>; L=$(grep -n "Restarted application" $S/run.log | tail -1 | cut -d: -f1); wc -l < $S/run.log; tail -n +$L $S/run.log | grep -c "didChangeMetrics\|EXCEPTION"
711
0
The errors have stopped. They came from my temporary debug code still being hooked into panels that were already open, and a hot restart cleared them. The log has been clean since the restart, and the TV app reopened where you were.
The rest is unchanged: nothing is committed yet, and the debug app is still running on the TV.
'~/Desktop/Screenshots/tv-130502.png'for the top bar buttons can you add a 50% opacity dash background color to make sure the texts are always reasdable
~/Desktop/Screenshots/tv-130502.png
[image]
I'm taking "dash" to mean a dark background. Finding the top-bar chip widget:
$ grep -n "'Ask'\|'Slideshow'" lib/screens/*.dart | head; grep -n "class _TopChip\|class _Chip\|class _ActionChip\|class .*Chip" -r lib | head
lib/screens/settings_screen.dart:135: _SectionHeader('Slideshow'),
lib/screens/slideshow_screen.dart:590: label: 'Ask',
lib/screens/viewer_screen.dart:404: label: 'Ask',
lib/screens/viewer_screen.dart:413: label: 'Slideshow',
lib/screens/settings_screen.dart:334:class _ChoiceChips<T> extends StatelessWidget {
lib/screens/slideshow_screen.dart:614:class _GlassChip extends StatefulWidget {
lib/screens/slideshow_screen.dart:630:class _GlassChipState extends State<_GlassChip> {
$ sed -n 395,425p lib/screens/viewer_screen.dart; sed -n 580,700p lib/screens/slideshow_screen.dart
fontSize: 13,
shadows: [Shadow(blurRadius: 8, color: Colors.black54)],
),
),
],
),
),
_TopPill(
icon: LucideIcons.sparkles,
label: 'Ask',
focusNode: _askFocus,
onActivate: _ask,
onDown: () => _imageFocus.requestFocus(),
onRight: () => _slideshowFocus.requestFocus(),
),
const SizedBox(width: 10),
_TopPill(
icon: LucideIcons.play,
label: 'Slideshow',
focusNode: _slideshowFocus,
onActivate: _startSlideshow,
onDown: () => _imageFocus.requestFocus(),
onLeft: () => _askFocus.requestFocus(),
),
],
),
),
),
// Bare chevron hints (fill on hover; D-pad uses Left/Right directly).
Positioned(
left: 18,
icon: _paused
? LucideIcons.play
: LucideIcons.pause,
label: _paused ? 'Play' : 'Pause',
onTap: _togglePause,
focusNode: _playFocus,
),
const SizedBox(width:
…[3146 more chars]
$ grep -n "class _PillButton" -A60 lib/screens/viewer_screen.dart | grep -n "color\|decoration\|border"
(Bash completed with no output)
$ L=$(grep -n "class _TopPill " lib/screens/viewer_screen.dart | cut -d: -f1); sed -n "$L,$((L+110))p" lib/screens/viewer_screen.dart | grep -n "color:\|alpha"
69: color: active ? Colors.white : Colors.transparent,
72: color: Colors.white.withValues(alpha: active ? 1 : 0.4)),
77: Icon(widget.icon, size: 16, color: fg),
81: color: fg,
$ python3 - <<'EOF'
import re
p='lib/screens/viewer_screen.dart'
s=open(p).read()
old="color: active ? Colors.white : Colors.transparent,"
assert s.count(old)==1
s=s.replace(old,"""// Dark when idle so the label reads over bright photos.
color: active
? Colors.white
: Colors.black.withValues(alpha: 0.5),""")
open(p,'w').write(s)
p='lib/screens/slideshow_screen.dart'
s=open(p).read()
old="color: _hover ? Colors.white : Colors.white.withValues(alpha: 0.14),"
assert s.count(old)==1
s=s.replace(old,"""// Dark when idle so the label reads over bright photos.
color: _hover ? Colors.white : Colors.black.withValues(alpha: 0.5),""")
open(p,'w').write(s)
EOF
S=<scratch>; echo r > $S/run.fifo; sleep 4; tail -1 $S/run.log; adb exec-out screencap -p > $S/s3.png; sips -Z 960 $S/s3.png >/dev/null
Reloaded 2 of 2219 libraries in 3,044ms (compile: 53 ms, reload: 1416 ms, reassemble: 1511 ms).
[This command modified 2 files you've previously read: lib/screens/slideshow_screen.dart, lib/screens/viewer_screen.dart. Call Read before editing.]
<scratch>/s3.png
[image]
The controls have auto-hidden. Pressing Up to bring them back:
$ S=<scratch>; adb shell input keyevent DPAD_UP; sleep 1; adb exec-out screencap -p > $S/s3.png; sips -c 200 960 --cropOffset 0 960 $S/s3.png --out $S/s3c.png >/dev/null
(Bash completed with no output)
<scratch>/s3c.png
[image]
I gave the top-bar buttons a 50% black background when they're not selected. The selected button stays solid white. I assumed "dash" was a typo for "dark". It's hot-reloaded on the TV. In the screenshot I took, "Ask" now reads clearly over the bright sky, where before it nearly disappeared.
I changed the same thing on the slideshow's top bar (the Play/Pause and Ask buttons) so the two screens match. Those were faint white at 14% before.
is ther a way I can record the screen of the TV? from this mac
I already have an alis to capture screenshot using tvshot
$ grep -n "tvshot" ~/.zshrc ~/.zprofile ~/.aliases 2>/dev/null; which scrcpy
Exit code 1
~/.zshrc:167:unalias tvshot 2>/dev/null
~/.zshrc:168:tvshot() {
scrcpy not found
$ sed -n 160,185p ~/.zshrc
# bun
export BUN_INSTALL="<scratch>/bun"
export PATH="$BUN_INSTALL/bin:$PATH"
# Screenshot the Sony TV over wireless adb. Finds the TV by model, connecting
# via wireless-debugging discovery if needed, so port changes don't matter.
unalias tvshot 2>/dev/null
tvshot() {
local s=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}')
if [ -z "$s" ]; then
local addr=$(adb mdns services | awk '/_adb-tls-connect/ {print $3; exit}')
[ -n "$addr" ] && adb connect "$addr" >/dev/null && s=$addr
fi
if [ -z "$s" ]; then
echo "TV not reachable — turn on Settings › System › Developer options › Wireless debugging on the TV."
return 1
fi
local f=~/Desktop/Screenshots/tv-$(date +%H%M%S).png
adb -s "$s" exec-out screencap -p > "$f" && echo "Saved $f"
}
# Added by DevRelay installer
export PATH="$HOME/.devrelay/bin:$PATH"
Yes. Android TV has a built-in screen recorder (adb shell screenrecord), so you don't need to install anything. I'll add a tvrec command next to your tvshot, in the same style:
{
"replace_all": false,
"file_path": "~/.zshrc",
"old_string": " adb -s \"$s\" exec-out screencap -p > \"$f\" && echo \"Saved $f\"\n}\n",
"new_string": " adb -s \"$s\" exec-out screencap -p > \"$f\" && echo \"Saved $f\"\n}\n\n# Record the Sony TV's screen: `tvrec`, press Enter to stop (Android caps a\n# recording at 3 minutes). Saves an .mp4 next to the screenshots.\nunalias tvrec 2>/dev/null\ntvrec() {\n local s=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}')\n if [ -z \"$s\" ]; then\n local addr=$(adb mdns services | awk '/_adb-tls-connect/ {print $3; exit}')\n [ -n \"$addr\" ] && adb connect \"$addr\" >/dev/null && s=$addr\n fi\n if [ -z \"$s\" ]; then\n echo \"TV not reachable — turn on Settings › System › Developer options › Wireless debugging on the TV.\"\n return 1\n fi\n local remote=/sdcard/tvrec.mp4\n local f=~/Desktop/Screenshots/tv-$(date +%H%M%S).mp4\n adb -s \"$s\" shell \"screenrecord --bit-rate 20000000 $remote\" &\n local pid=$!\n echo \"Recording… press Enter to stop.\"\n read -r\n # SIGINT lets screenrecord finish the file properly.\n adb -s \"$s\" shell pkill -INT screenrecord\n wait $pid 2>/dev/null\n sleep 1\n adb -s \"$s\" pull $remote \"$f\" >/dev/null && adb -s \"$s\" shell rm $remote &&\n echo \"Saved $f\"\n}\n"
}
The file ~/.zshrc has been updated successfully.
$ S=<scratch>; ( sleep 6; echo ) | zsh -ic 'tvrec' 2>&1 | tail -3; f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); ls -la "$f"; ffprobe -v error -show_entries format=duration:stream=width,height,codec_name -of compact "$f" 2>/dev/null || mdls -name kMDItemDurationSeconds -name kMDItemPixelWidth "$f"
Recording… press Enter to stop.
/sdcard/tvrec.mp4: 1 file pulled, 0 skipped. 8.3 MB/s (127411 bytes in 0.015s)
Saved ~/Desktop/Screenshots/tv-130718.mp4
-rw-r--r--@ 1 souvikbiswas staff 127411 Oct 4 13:07 ~/Desktop/Screenshots/tv-130718.mp4
kMDItemDurationSeconds = (null)
kMDItemPixelWidth = (null)
A test recording saved. I'm checking that the file is a complete, playable video:
$ f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); S=<scratch>; cat > $S/d.swift <<EOF
import AVFoundation
let a = AVURLAsset(url: URL(fileURLWithPath: "$f"))
let sem = DispatchSemaphore(value: 0)
Task { let d = try await a.load(.duration); let t = try await a.loadTracks(withMediaType: .video); let sz = try await t.first!.load(.naturalSize); print("duration \(d.seconds)s size \(sz)"); sem.signal() }
sem.wait()
EOF
swift $S/d.swift 2>&1 | tail -2
[#NoUseUnstructuredThrowingTask]: <https://docs.swift.org/compiler/documentation/diagnostics/no-use-throwing-unstructured-task>
duration 0.0s size (1920.0, 1080.0)
The file is 1080p but only 0 s long. The TV screen didn't change during the test, and Android's recorder only saves frames when something on screen moves. I'll test again while moving through photos:
$ S=<scratch>; rm -f ~/Desktop/Screenshots/tv-130718.mp4; ( sleep 2; adb shell input keyevent DPAD_DOWN; sleep 2; adb shell input keyevent DPAD_RIGHT; sleep 2; adb shell input keyevent DPAD_LEFT; sleep 2; echo ) | zsh -ic 'tvrec' 2>&1 | tail -1; f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); sed -i '' "s|URL(fileURLWithPath: \".*\")|URL(fileURLWithPath: \"$f\")|" $S/d.swift; swift $S/d.swift 2>&1 | tail -1
Saved ~/Desktop/Screenshots/tv-130753.mp4
duration 6.8034888888888885s size (1920.0, 1080.0)
Yes. I added a tvrec command to your ~/.zshrc, right below tvshot, and it works.
tvrec # starts recording; press Enter to stop
It finds the TV the same way tvshot does and saves ~/Desktop/Screenshots/tv-HHMMSS.mp4 at 1080p. It uses the recorder built into Android TV, so nothing extra is installed. Open a new terminal, or run source ~/.zshrc, to get it.
-
Tested: I recorded while pressing the remote buttons over adb and got a 6.8 s, 1920×1080 video. The test clip is at
~/Desktop/Screenshots/tv-130753.mp4if you want to look at it or delete it. - Limits: a single recording stops at 3 minutes, and there's no sound. A screen that doesn't change records as almost nothing; my first test on a still photo came out 0 s long.
-
If you need sound or longer clips (for example the voice reading answers aloud in your demo):
brew install scrcpy, thenscrcpy --record demo.mp4. It mirrors the TV in a window on the Mac and records audio too, on Android 11 and later. I haven't installed it. Tell me if you want it.
can it record voice that I speak during the input and the output voice too?
The TV's own sound (Gemma's answers read aloud) can be recorded. Your voice is harder, because it goes into the remote's mic and only the TV's voice search hears it.
tvrec (now) |
scrcpy | Mac microphone | |
|---|---|---|---|
| Video | ✅ | ✅ | — |
| TV sound / spoken answers | ❌ | ✅ likely | ✅ if the Mac is in the room |
| Your voice | ❌ | ⚠️ probably not | ✅ |
- TV sound: scrcpy can grab everything the TV plays, which would include the spoken answers. I haven't tested it with this TV's text-to-speech yet. That's the first thing to check.
-
Your voice through scrcpy: it can record from the TV's microphone (
--audio-source=mic). But while you hold the mic button, Google's voice typing is using that mic, and Android usually gives other recorders silence then. It's also unclear whether a Bluetooth remote's mic counts as the TV's mic at all. I expect this not to work. -
Most reliable for a demo: scrcpy records the TV's picture and sound while the Mac's microphone records the room. I'd make a
tvrecthat runs both and merges them into one .mp4 with ffmpeg. Your voice comes from the Mac mic, and the answers come straight from the TV so they're clean.
The one thing in your voice that always ends up on the TV screen is the text: voice typing writes your question into the box.
Want me to install scrcpy and ffmpeg (brew install scrcpy ffmpeg), check that the answers are recorded, and make tvrec record both? You'd need to say a test question into the remote while I record.
yes, set it up
$ brew install scrcpy ffmpeg 2>&1 | tail -3; scrcpy --version | head -1; ffmpeg -version | head -1; ffmpeg -f avfoundation -list_devices true -i "" 2>&1 | grep -A10 "audio devices"
You can install adb from Homebrew Cask:
brew install --cask android-platform-tools
scrcpy 4.1 <https://github.com/Genymobile/scrcpy>
ffmpeg version 9.0.2 Copyright (c) 2000-2026 the FFmpeg developers
[AVFoundation indev @ 0x7b21018140] AVFoundation audio devices:
[AVFoundation indev @ 0x7b21018140] [0] ZoomAudioDevice
[AVFoundation indev @ 0x7b21018140] [1] MacBook Pro Microphone
[in#0 @ 0x7b21018000] Error opening input: Input/output error
Error opening input file .
Error opening input files: Input/output error
$ S=<scratch>; SER=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}'); adb exec-out screencap -p > $S/a.png; sips -Z 640 $S/a.png >/dev/null; echo $SER
<tv-adb-serial>
<scratch>/a.png
[image]
$ S=<scratch>; adb shell input keyevent DPAD_UP; sleep 0.5; adb shell input keyevent DPAD_CENTER; sleep 2; adb shell input keyevent BACK; sleep 1; adb exec-out screencap -p > $S/a.png; sips -Z 640 $S/a.png >/dev/null
(Bash completed with no output)
<scratch>/a.png
[image]
$ S=<scratch>; adb shell input keyevent DPAD_LEFT; sleep 0.5; adb shell input keyevent DPAD_CENTER; sleep 2.5; adb shell input keyevent BACK; sleep 1; adb exec-out screencap -p > $S/a.png; sips -Z 640 $S/a.png >/dev/null
(Bash completed with no output)
<scratch>/a.png
[image]
$ S=<scratch>; adb shell input keyevent DPAD_DOWN; sleep 3; adb exec-out screencap -p > $S/a.png; sips -Z 640 $S/a.png >/dev/null; adb shell dumpsys input_method | grep mInputShown
mInputShown=true
<scratch>/a.png
[image]
Ask Gemma is open on the TV. Next I start a scrcpy recording, have the TV ask "What is this?", and check whether the spoken answer was captured:
$ S=<scratch>; SER=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}'); rm -f $S/t1.mkv; scrcpy -s "$SER" --no-window --no-control --audio-source=playback --audio-dup --record=$S/t1.mkv --time-limit=22 > $S/scrcpy.log 2>&1 &
sleep 3; adb shell input keyevent BACK; sleep 1; adb shell input keyevent DPAD_DOWN; sleep 0.7; adb shell input keyevent DPAD_CENTER; wait; tail -5 $S/scrcpy.log; ffmpeg -i $S/t1.mkv -af volumedetect -vn -f null - 2>&1 | grep -E "Stream|mean_volume|max_volume"; ffmpeg -i $S/t1.mkv -af silencedetect=n=-45dB:d=1 -vn -f null - 2>&1 | grep silence_ | head
INFO: ADB device found:
INFO: --> (tcpip) <tv-adb-serial> device BRAVIA_4K_AE2
INFO: Recording started to matroska file: <scratch>/t1.mkv
INFO: Time limit reached
INFO: Recording complete to matroska file: <scratch>/t1.mkv
Stream #0:0: Video: h264 (High), yuv420p(tv, bt709, progressive), 1920x1080, 1k tbr, 1k tbn
Stream #0:1: Audio: opus, 48000 Hz, stereo, fltp, start 0.089000
Stream mapping:
Stream #0:1 -> #0:0 (opus (native) -> pcm_s16le (native))
Stream #0:0: Audio: pcm_s16le, 48000 Hz, stereo, s16, 1536 kb/s
[Parsed_volumedetect_0 @ 0x75c7024a80] mean_volume: -22.3 dB
[Parsed_volumedetect_0 @ 0x75c7024a80] max_volume: -4.0 dB
[Parsed_silencedetect_0 @ 0x7c61024a80] silence_start: 0.0595
[Parsed_silencedetect_0 @ 0x7c61024a80] silence_end: 8.995917 | silence_duration: 8.936417
Sound was captured: silence for the first 9 s, then about 13 s of audio. I'll check a frame to confirm it's the answer being read:
$ S=<scratch>; ffmpeg -y -v error -ss 12 -i $S/t1.mkv -frames:v 1 -vf scale=640:-1 $S/f.png
(Bash completed with no output)
<scratch>/f.png
File does not exist. Note: your current working directory is ~/MobileDev/Flutter/tv_gallery_gemma.
$ S=<scratch>; ffmpeg -y -v error -i $S/t1.mkv -vf "select=gte(t\,12),scale=640:-1" -frames:v 1 $S/f.png; ls $S/f.png; ffprobe -v error -show_entries format=duration -of csv=p=0 $S/t1.mkv
ls: <scratch>/f.png: No such file or directory
21.748000
<scratch>/f.png
File does not exist. Note: your current working directory is ~/MobileDev/Flutter/tv_gallery_gemma.
$ S=<scratch>; ffprobe -v error -select_streams v -show_entries packet=pts_time -of csv=p=0 $S/t1.mkv | tail -3; ffmpeg -y -v error -sseof -1 -i $S/t1.mkv -update 1 -vf scale=640:-1 $S/f.png; ls $S/f.png
Exit code 1
8.674000
8.774000
8.874000
ls: <scratch>/f.png: No such file or directory
<scratch>/f.png
File does not exist. Note: your current working directory is ~/MobileDev/Flutter/tv_gallery_gemma.
$ S=<scratch>; ffmpeg -y -v error -i $S/t1.mkv -an -update 1 -vf scale=640:-1 $S/f.png && ls $S/f.png
<scratch>/f.png
<scratch>/f.png
[image]
$ python3 - <<'EOF'
import os,re
p=os.path.expanduser('~/.zshrc')
s=open(p).read()
a=s.index("# Record the Sony TV's screen")
b=s.index("}\n",s.index("tvrec() {"))+2
new='''# Record the Sony TV for a demo: `tvrec`, press Enter to stop. Captures the
# TV's picture and sound (e.g. answers read aloud) with scrcpy, plus your
# voice through the Mac's microphone, mixed into one .mp4 next to the
# screenshots. `tvrec --no-mic` skips the microphone.
unalias tvrec 2>/dev/null
tvrec() {
local s=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}')
if [ -z "$s" ]; then
local addr=$(adb mdns services | awk '/_adb-tls-connect/ {print $3; exit}')
[ -n "$addr" ] && adb connect "$addr" >/dev/null && s=$addr
fi
if [ -z "$s" ]; then
echo "TV not reachable — turn on Settings › System › Developer options › Wireless debugging on the TV."
return 1
fi
local mic=1; [ "$1" = "--no-mic" ] && mic=0
local tmp=$(mktemp -d)
local f=~/Desktop/Screenshots/tv-$(date +%H%M%S).mp4
# --audio-dup keeps the sound playing on the TV while it's recorded.
scrcpy -s "$s" --no-window --no-control --audio-source=playback --audio-dup \\
--record="$tmp/tv.mkv" >"$tmp/scrcpy.log" 2>&1 &
local tv=$!
local m
if [ $mic = 1 ]; then
ffmpeg -nostdin -loglevel error -f avfoundation -i ":MacBook Pro Microphone" \\
-c:a aac -b:a 160k "$tmp/mic.m4a" &
m=$!
fi
echo "Recording… press Enter to stop."
read -r
# SIGINT lets both finish their files properly.
kill -INT $tv ${m:-} 2>/dev/null
wait $tv ${m:-} 2>/dev/null
if [ ! -s "$tmp/tv.mkv" ]; then
echo "Recording failed:"; cat "$tmp/scrcpy.log"; return 1
fi
echo "Saving…"
# scrcpy only sends frames when the screen changes; hold the last frame so
# the video runs as long as the sound.
local v="tpad=stop=-1:stop_mode=clone,fps=30,format=yuv420p"
if [ $mic = 1 ] && [ -s "$tmp/mic.m4a" ]; then
ffmpeg -loglevel error -i "$tmp/tv.mkv" -i "$tmp/mic.m4a" -filter_complex \\
"[0:v]$v[v];[0:a][1:a]amix=inputs=2:duration=longest:normalize=0[a]" \\
-map "[v]" -map "[a]" -shortest -c:v libx264 -crf 18 -preset veryfast \\
-c:a aac -b:a 192k -movflags +faststart "$f"
else
ffmpeg -loglevel error -i "$tmp/tv.mkv" -vf "$v" -shortest -c:v libx264 \\
-crf 18 -preset veryfast -c:a aac -b:a 192k -movflags +faststart "$f"
fi && rm -rf "$tmp" && echo "Saved $f"
}
'''
s=s[:a]+new+s[b:]
open(p,'w').write(s)
EOF
zsh -n ~/.zshrc && echo syntax ok; grep -n "^tvrec\|^tvshot" ~/.zshrc
syntax ok
168:tvshot() {
187:tvrec() {
Syntax checks out. Before handing it over, I'll test the merge step on the TV recording I already have, plus a stand-in mic track:
$ S=<scratch>; cd $S; ffmpeg -y -loglevel error -f lavfi -i "sine=f=440:d=20" -c:a aac mic.m4a; v="tpad=stop=-1:stop_mode=clone,fps=30,format=yuv420p"; ffmpeg -y -loglevel error -i t1.mkv -i mic.m4a -filter_complex "[0:v]$v[v];[0:a][1:a]amix=inputs=2:duration=longest:normalize=0[a]" -map "[v]" -map "[a]" -shortest -c:v libx264 -crf 18 -preset veryfast -c:a aac -b:a 192k -movflags +faststart out.mp4 && ffprobe -v error -show_entries stream=codec_type,duration -of compact out.mp4; ffmpeg -y -loglevel error -i t1.mkv -vf "$v" -shortest -c:v libx264 -crf 18 -preset veryfast -c:a aac -movflags +faststart out2.mp4 && ffprobe -v error -show_entries stream=codec_type,duration -of compact out2.mp4
Exit code 1
(eval):1: bad math expression: ':' without '?'
That caught a real bug: zsh reads $v[v] as an array lookup. I'm fixing it in tvrec and rerunning:
$ sed -i '' 's/"\[0:v\]\$v\[v\];/"[0:v]${v}[v];/' ~/.zshrc && grep -n '\[0:v\]' ~/.zshrc; S=<scratch>; cd $S; v="tpad=stop=-1:stop_mode=clone,fps=30,format=yuv420p"; ffmpeg -y -loglevel error -i t1.mkv -i mic.m4a -filter_complex "[0:v]${v}[v];[0:a][1:a]amix=inputs=2:duration=longest:normalize=0[a]" -map "[v]" -map "[a]" -shortest -c:v libx264 -crf 18 -preset veryfast -c:a aac -b:a 192k -movflags +faststart out.mp4 && ffprobe -v error -show_entries stream=codec_type,duration -of compact out.mp4; ffmpeg -y -loglevel error -i t1.mkv -vf "$v" -shortest -c:v libx264 -crf 18 -preset veryfast -c:a aac -movflags +faststart out2.mp4 && ffprobe -v error -show_entries stream=codec_type,duration -of compact out2.mp4
224: "[0:v]${v}[v];[0:a][1:a]amix=inputs=2:duration=longest:normalize=0[a]" \
stream|codec_type=video|duration=21.766667
stream|codec_type=audio|duration=21.693333
stream|codec_type=video|duration=22.900000
stream|codec_type=audio|duration=21.693333
Shell cwd was reset to ~/MobileDev/Flutter/tv_gallery_gemma
tvrec is set up. Open a new terminal (or run source ~/.zshrc) to pick it up.
tvrec # TV picture + TV sound + your voice via the Mac mic → press Enter to stop
tvrec --no-mic # TV picture and sound only
It saves ~/Desktop/Screenshots/tv-HHMMSS.mp4. While it records, the TV keeps playing sound normally, and there's no 3-minute limit any more.
What I tested:
- TV sound: recording a real question on the TV captured Gemma's spoken answer about the Shanghai photo, clearly, starting at 9 s.
- Merge step: I ran it on that recording plus a fake mic track and got a 21.7 s mp4 with the picture and sound lined up. This test caught a zsh quoting bug in the merge command, which I fixed.
- Not tested: a full run with your real voice through the Mac mic, since you're testing that yourself.
On your first run:
- Mic permission: macOS will ask whether your terminal app can use the microphone. Allow it, or your voice will be silent.
- Echo: the Mac mic also hears the TV's speakers, so the spoken answers may echo slightly over the clean copy. If that bothers you, I can turn the mic down while answers are playing, or switch to an external mic. Your mic list shows only the MacBook's built-in mic and a Zoom virtual device right now.
$ S=<scratch>; grep -n -B6 -A8 "dead thread" $S/run.log | tail -20
1587-D/MediaCodec( 7444): flushMediametrics
1588-D/MediaCodec( 7444): notifyMediaStatistics hash (0)
1589-D/MediaCodec( 7444): flushMediametrics
1590-W/MediaCodec( 7444): no metrics handle found
1591-I/hw-BpHwBinder( 7444): onLastStrongRef automatically unlinking death recipients
1592-I/ExoPlayerImpl( 7444): Release 5c4814 [AndroidXMedia3/1.4.1] [BRAVIA_AE2, BRAVIA 4K AE2, Sony, 34] [media3.common, media3.exoplayer, media3.decoder, media3.datasource, media3.extractor]
1593:W/MessageQueue( 7444): Handler (android.os.Handler) {3ef5f55} sending message to a Handler on a dead thread
1594:W/MessageQueue( 7444): java.lang.IllegalStateException: Handler (android.os.Handler) {3ef5f55} sending message to a Handler on a dead thread
1595-W/MessageQueue( 7444): at android.os.MessageQueue.enqueueMessage(MessageQueue.java:560)
1596-W/MessageQueue( 7444): at android.os.Handler.enqueueMessage(Handler.java:786)
1597-W/MessageQueue( 7444): at android.os.Handler.sendMessageAtTime(Handler.java:735)
1598-W/MessageQueue( 7444): at android.os.Handler.sendMessageDelayed(Handler.java:705)
1599-W/MessageQueue( 7444): at android.os.Handler.sendMessage(Handler.java:643)
1600-W/MessageQueue( 7444): at android.os.Message.sendToTarget(Message.java:468)
1601-W/MessageQueue( 7444): at androidx.media3.common.util.SystemHandlerWrapper$SystemMessage.sendToTarget(SystemHandlerWrapper.java:160)
1602-W/MessageQueue( 7444): at androidx.media3.exoplayer.ExoPlayerImplInternal.stop(ExoPlayerImplInternal.java:3
…[3 more chars]
The new log alert is harmless. It's an Android warning from the video player as it shuts down: after a video was released, its stop() call reached a background thread that had already ended. The app didn't crash, and it isn't related to the Gemma changes. The video player code hasn't changed.
got this:
<pasted_content id="396a">
~ » tvrec me@mac
[2] 86157
[4] 86158
Recording… press Enter to stop.
[2] 86157 done scrcpy -s "$s" --no-window --no-control --audio-source=playback --audio-dup
[4] + 86158 exit 255 ffmpeg -nostdin -loglevel error -f avfoundation -i ":MacBook Pro Microphone"
Recording failed:
/opt/homebrew/Cellar/scrcpy/4.1_1/share/scrcpy/scrcpy-server: 1 file pushed, 0 skipped. 81.4 MB/s (733706 bytes in 0.009s)
[server] INFO: Device: [Sony] Sony BRAVIA 4K AE2 (Android 14)
scrcpy 4.1 https://github.com/Genymobile/scrcpy
INFO: ADB device found:
INFO: --> (tcpip) <tv-adb-serial> device BRAVIA_4K_AE2
</pasted_content id="396a">
Two separate failures here: the microphone recorder exited straight away (code 255), and scrcpy stopped before it ever began writing the video. I'll reproduce each on its own:
$ S=<scratch>; cd $S; timeout 4 ffmpeg -y -nostdin -f avfoundation -i ":MacBook Pro Microphone" -t 2 m.m4a 2>&1 | tail -4; echo "exit $?"; timeout 4 ffmpeg -y -nostdin -f avfoundation -i ":1" -t 2 m.m4a 2>&1 | tail -3
encoder : Lavc63.1.102 aac
size= 0KiB time=00:00:00.44 bitrate= 0.8kbits/s speed=0.895x elapsed=0:00:00.50
size= 0KiB time=00:00:00.96 bitrate= 0.4kbits/s speed=0.954x elapsed=0:00:01.00
size= 0KiB time=00:00:01.46 bitrate= 0.2kbits/s speed=0.967x elapsed=0:00:01.51
[out#0/ipod @ 0x7c6511c840] video:0KiB audio:16KiB subtitle:0KiB other streams:0KiB global headers:0KiB muxing overhead: 9.286878%
size= 17KiB time=00:00:02.00 bitrate= 69.8kbits/s speed=0.987x elapsed=0:00:02.02
[aac @ 0x7c648b8a80] Qavg: 128.304
exit 0
size= 0KiB time=00:00:00.44 bitrate= 0.8kbits/s speed=0.89x elapsed=0:00:00.50
size= 0KiB time=00:00:00.96 bitrate= 0.4kbits/s speed=0.952x elapsed=0:00:01.00
size= 0KiB time=00:00:01.46 bitrate= 0.2kbits/s speed=0.969x elapsed=0:00:01.50
size= 0KiB time=00:00:01.94 bitrate= 0.2kbits/s speed=0.964x elapsed=0:00:02.01
[out#0/ipod @ 0x77690ac840] video:0KiB audio:15KiB subtitle:0KiB other streams:0KiB global headers:0KiB muxing overhead: 9.242797%
size= 17KiB time=00:00:02.00 bitrate= 69.3kbits/s speed=0.987x elapsed=0:00:02.02
[aac @ 0x7768944a80] Qavg: 107.855
Shell cwd was reset to ~/MobileDev/Flutter/tv_gallery_gemma
$ ls -lat $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/ | grep "tmp\." | head -3; d=$(ls -dt $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/tmp.* | head -1); ls -la $d; cat $d/scrcpy.log
drwx------@ 4 souvikbiswas staff 128 Oct 4 13:19 tmp.Rn3ryW7fvX
total 816
drwx------@ 4 souvikbiswas staff 128 Oct 4 13:19 .
drwx------@ 377 souvikbiswas staff 12064 Oct 4 13:20 ..
-rw-r--r--@ 1 souvikbiswas staff 352810 Oct 4 13:20 mic.m4a
-rw-r--r--@ 1 souvikbiswas staff 376 Oct 4 13:20 scrcpy.log
/opt/homebrew/Cellar/scrcpy/4.1_1/share/scrcpy/scrcpy-server: 1 file pushed, 0 skipped. 81.4 MB/s (733706 bytes in 0.009s)
[server] INFO: Device: [Sony] Sony BRAVIA 4K AE2 (Android 14)
scrcpy 4.1 <https://github.com/Genymobile/scrcpy>
INFO: ADB device found:
INFO: --> (tcpip) <tv-adb-serial> device BRAVIA_4K_AE2
The mic part worked: your run saved about 18 s of mic audio. The problem is scrcpy. It connected to the TV and then never started recording. Reproducing that in an interactive shell like yours:
$ ( sleep 10; echo ) | zsh -ic 'tvrec --no-mic' 2>&1 | tail -8
Command did not complete within its 90s timeout and was moved to the background (ID: bt2a7o097). Output is being written to: <tmp> You will be notified when it completes. To check interim output, use Read on that file path.
$ pgrep -fl "scrcpy|ffmpeg" ; d=$(ls -dt $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/tmp.* | head -1); ls -la $d; cat $d/scrcpy.log; cat <tmp>
87352 scrcpy -s <tv-adb-serial> --no-window --no-control --audio-source=playback --audio-dup --record=<tmp>
87357 adb -s <tv-adb-serial> shell CLASSPATH=/data/local/tmp/scrcpy-server.jar app_process / com.genymobile.scrcpy.Server 4.1 scid=5234291b log_level=info audio_source=playback audio_dup=true control=false
total 8712
drwx------@ 4 souvikbiswas staff 128 Oct 4 13:21 .
drwx------@ 378 souvikbiswas staff 12096 Oct 4 13:21 ..
-rw-r--r--@ 1 souvikbiswas staff 185 Oct 4 13:21 scrcpy.log
-rw-r--r--@ 1 souvikbiswas staff 4194304 Oct 4 13:21 tv.mkv
/opt/homebrew/Cellar/scrcpy/4.1_1/share/scrcpy/scrcpy-server: 1 file pushed, 0 skipped. 93.6 MB/s (733706 bytes in 0.007s)
[server] INFO: Device: [Sony] Sony BRAVIA 4K AE2 (Android 14)
$ kill -TERM 87352; sleep 3; pgrep -fl scrcpy; cat <tmp> | tail -5; d=$(ls -dt $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/tmp.* | head -1); cat $d/scrcpy.log 2>/dev/null | tail -3; ls ~/Desktop/Screenshots/*.mp4 | tail -2
scrcpy 4.1 <https://github.com/Genymobile/scrcpy>
INFO: ADB device found:
INFO: --> (tcpip) <tv-adb-serial> device BRAVIA_4K_AE2
~/Desktop/Screenshots/tv-131849.mp4
~/Desktop/Screenshots/tv-132112.mp4
$ python3 - <<'EOF'
import os
p=os.path.expanduser('~/.zshrc')
s=open(p).read()
def r(a,b):
global s
assert a in s, a
s=s.replace(a,b)
r('''# Record the Sony TV for a demo: `tvrec`, press Enter to stop. Captures the
# TV's picture and sound (e.g. answers read aloud) with scrcpy, plus your
# voice through the Mac's microphone, mixed into one .mp4 next to the
# screenshots. `tvrec --no-mic` skips the microphone.''','''# Record the Sony TV for a demo: `tvrec`, press Enter to stop. Captures the
# TV's picture and sound (e.g. answers read aloud) with scrcpy, plus your
# voice through the Mac's microphone, mixed into one .mp4 next to the
# screenshots. The TV's speakers go quiet while recording so the mic doesn't
# pick them up as an echo; the sound is still recorded directly.
# `tvrec --no-mic` skips the microphone and keeps the TV's sound on.''')
r(''' # --audio-dup keeps the sound playing on the TV while it's recorded.
scrcpy -s "$s" --no-window --no-control --audio-source=playback --audio-dup \\
--record="$tmp/tv.mkv" >"$tmp/scrcpy.log" 2>&1 &
local tv=$!''',''' # Capturing the TV's sound silences its speakers; --audio-dup keeps them
# playing (only wanted without the mic, which would hear them).
local dup=(); [ $mic = 0 ] && dup=(--audio-dup)
scrcpy -s "$s" --no-window --no-control --audio-source=playback $dup \\
--record="$tmp/tv.mkv" >"$tmp/scrcpy.log" 2>&1 &
local tv=$!
# Wait until scrcpy is actually writing before saying we're recording.
local i=0
while [ ! -s "$tmp/tv.mkv" ] && kill -0 $tv 2>/dev/null && [ $i -lt 100 ]; do
sleep 0.1; i=$((i+1))
done
if [ ! -s "$tmp/tv.mkv" ]; then
kill -TERM $tv 2>/dev/null
echo "Couldn't start recording the TV:"; cat "$tmp/scrcpy.log"; return 1
fi''')
r(''' echo "Recording… press Enter to stop."
read -r
# SIGINT lets both finish their files properly.
kill -INT $tv ${m:-} 2>/dev/null''',''' echo "Recording… press Enter to stop."
read -r
# SIGTERM lets both finish their files properly (scrcpy can ignore SIGINT
# when run as a background job).
kill -TERM $tv ${m:-} 2>/dev/null''')
r(''' if [ ! -s "$tmp/tv.mkv" ]; then
echo "Recording failed:"; cat "$tmp/scrcpy.log"; return 1
fi
echo "Saving…"''',''' echo "Saving…"''')
open(p,'w').write(s)
EOF
zsh -n ~/.zshrc && echo ok
ok
Testing the full path (TV plus mic) the same way your shell runs it:
$ ( sleep 8; echo ) | zsh -ic 'tvrec' 2>&1 | grep -v "^\[" | tail -4; f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); ffprobe -v error -show_entries stream=codec_type,duration -of compact "$f"
Couldn't start recording the TV:
/opt/homebrew/Cellar/scrcpy/4.1_1/share/scrcpy/scrcpy-server: 1 file pushed, 0 skipped. 92.2 MB/s (733706 bytes in 0.008s)
stream|codec_type=video|duration=108.933333
stream|codec_type=audio|duration=108.953000
That reproduces your failure. I suspect scrcpy only creates the file once the TV screen changes, and the screen was still both times. To check, I'll start a recording, wait, then press a key:
$ S=<scratch>; SER=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}'); rm -f $S/t2.mkv; scrcpy -s "$SER" --no-window --no-control --audio-source=playback --record=$S/t2.mkv > $S/s2.log 2>&1 & P=$!; for t in 2 4 6; do sleep 2; echo "t=$t size=$(stat -f %z $S/t2.mkv 2>/dev/null || echo none)"; done; adb shell input keyevent DPAD_UP; for t in 7 8 10; do sleep 1; echo "after key t=$t size=$(stat -f %z $S/t2.mkv 2>/dev/null || echo none)"; done; kill -TERM $P; wait $P; ffprobe -v error -show_entries stream=codec_type -of csv=p=0 $S/t2.mkv | tr '\n' ' '
t=2 size=0
t=4 size=0
t=6 size=0
after key t=7 size=262144
after key t=8 size=262144
after key t=10 size=262144
video audio
$ python3 - <<'EOF'
import os
p=os.path.expanduser('~/.zshrc')
s=open(p).read()
old=''' # Wait until scrcpy is actually writing before saying we're recording.
local i=0
while [ ! -s "$tmp/tv.mkv" ] && kill -0 $tv 2>/dev/null && [ $i -lt 100 ]; do
sleep 0.1; i=$((i+1))
done'''
assert old in s
s=s.replace(old,''' # scrcpy only starts writing once something on the TV screen changes; start
# the mic then too, so the two line up.
echo "Press any button on the TV remote to start…"
local i=0
while [ ! -s "$tmp/tv.mkv" ] && kill -0 $tv 2>/dev/null && [ $i -lt 600 ]; do
sleep 0.1; i=$((i+1))
done''')
open(p,'w').write(s)
EOF
zsh -n ~/.zshrc && echo ok; S=<scratch>; ( sleep 4; adb shell input keyevent DPAD_UP; sleep 6; echo ) | zsh -ic 'tvrec' 2>&1 | grep -v "^\[" | tail -4; f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); echo $f; ffprobe -v error -show_entries stream=codec_type,duration -of compact "$f"
ok
Press any button on the TV remote to start…
Recording… press Enter to stop.
Saving…
Saved ~/Desktop/Screenshots/tv-132426.mp4
~/Desktop/Screenshots/tv-132426.mp4
stream|codec_type=video|duration=7.633333
stream|codec_type=audio|duration=7.665000
$ rm -f ~/Desktop/Screenshots/tv-132112.mp4 ~/Desktop/Screenshots/tv-132426.mp4; rm -rf $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/tmp.Rn3ryW7fvX $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/tmp.9ovndpgV5O; ls ~/Desktop/Screenshots/*.mp4
~/Desktop/Screenshots/Voice Mode AI Agent.mp4
~/Desktop/Screenshots/tv-130753.mp4
~/Desktop/Screenshots/tv-131849.mp4
tvrec works now. Open a new terminal (or run source ~/.zshrc), then:
~ » tvrec
Press any button on the TV remote to start…
Recording… press Enter to stop.
Saving…
Saved ~/Desktop/Screenshots/tv-HHMMSS.mp4
Why your run failed: scrcpy doesn't start writing until something on the TV screen changes. Your screen was still, so it never started, and when you pressed Enter it quit with nothing saved. tvrec now waits for that first change, and pressing any button on the remote kicks it off. It also starts the mic at that moment, so your voice and the TV stay in sync. Separately, it now stops scrcpy with a signal the background process can't ignore; in my test it had ignored Ctrl-C and run on.
Echo: during tvrec the TV's speakers go quiet, so the Mac mic hears only you. Gemma's spoken answer still goes straight into the video, cleanly. You won't hear the answer in the room while recording, but it's in the file. tvrec --no-mic leaves the TV's speakers on.
Tested: I ran it the way your terminal does, with a simulated button press, and got a 7.6 s mp4 with the picture and the mixed audio lined up. I haven't been able to confirm by ear that the TV actually goes quiet, so check that on your first take.
I deleted my test recordings and the leftover temp files from your failed run. tv-131849.mp4 isn't one of mine as far as I can tell, so I left it. tv-130753.mp4 is my first silent test clip, so delete it if you don't want it.
there's a weird bug seems like, in the latest one there's a sound from the tv recorded even before the actual tv sound started, and it's overlapped with the actual tv sound
also no need to do fancy stuff for cancelling tv sound separwately, I have turned off tv sound while recording
$ ls -lt ~/Desktop/Screenshots/*.mp4 | head -3; f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); ffprobe -v error -show_entries stream=codec_type,start_time,duration -of compact "$f"; ls -dt $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/tmp.* | head -2
-rw-r--r--@ 1 souvikbiswas staff 1900247 Oct 4 13:25 ~/Desktop/Screenshots/tv-132510.mp4
-rw-r--r--@ 1 souvikbiswas staff 530955 Oct 4 13:18 ~/Desktop/Screenshots/tv-131849.mp4
-rw-r--r--@ 1 souvikbiswas staff 531157 Oct 4 13:08 ~/Desktop/Screenshots/tv-130753.mp4
stream|codec_type=video|start_time=0.000000|duration=25.266667
stream|codec_type=audio|start_time=0.056667|duration=25.234333
<tmp>
$ f=~/Desktop/Screenshots/tv-132510.mp4; ffmpeg -v error -i "$f" -af "astats=metadata=1:reset=25,ametadata=print:key=lavfi.astats.Overall.RMS_level" -vn -f null - 2>&1 | awk '/pts_time/ {split($0,a,"pts_time:"); t=a[2]} /RMS_level/ {split($0,b,"="); printf "%5.1fs %6.1f dB\n", t, b[2]}' | awk 'NR%2==1' | head -60
(Bash completed with no output)
$ f=~/Desktop/Screenshots/tv-132510.mp4; ffmpeg -v error -i "$f" -vn -af "asetnsamples=24000,astats=metadata=1:reset=1,ametadata=print:key=lavfi.astats.Overall.RMS_level:file=-" -f null - 2>/dev/null | paste - - | awk '{split($2,a,":"); split($NF,b,"="); printf "%s %s\n", a[2], b[2]}' | head -60
2720 -60.679237
25352 -59.313539
48968 -60.036571
73496 -44.973808
97016 -42.570723
121640 -46.826949
145160 -54.462430
169784 -59.963995
193256 -60.107647
216968 -60.166600
241400 -60.252781
264968 -60.595553
289544 -57.373680
313064 -42.247397
337880 -40.367841
361256 -42.251075
384968 -42.542862
409376 -39.869026
433016 -43.901877
457496 -58.547904
481016 -40.801929
505640 -39.061137
529208 -42.182243
553832 -48.722109
577352 -45.456716
600968 -18.391808
625448 -16.381484
649016 -18.839625
673592 -18.792132
697112 -17.496528
721736 -18.207228
745256 -29.211740
768920 -24.429086
793352 -16.417064
816920 -16.961133
841544 -19.864309
865064 -27.600557
889688 -22.012297
913208 -94.214709
937832 -20.262983
961304 -18.922062
985016 -19.438950
1009448 -25.601466
1032944 -20.620395
1057592 -20.690211
1081160 -21.930093
1105784 -35.974297
1129256 -94.770664
1152872 -inf
1177352 -inf
1200872 -inf
I'll line up when the answer appears on screen with when the loud audio starts:
$ S=<scratch>; f=~/Desktop/Screenshots/tv-132510.mp4; mkdir -p $S/fr; rm -f $S/fr/*; ffmpeg -v error -i "$f" -vf "fps=1,scale=480:-1,tile=5x5" -frames:v 1 $S/fr/tile.png; ls $S/fr
(eval):1: no matches found: <scratch>/fr/*
tile.png
<scratch>/fr/tile.png
[image]
$ S=<scratch>; ffprobe -v error -show_entries stream=codec_type,start_time,duration -of compact $S/t2.mkv; ffprobe -v error -select_streams v -show_entries packet=pts_time -of csv=p=0 $S/t2.mkv | head -2; ffprobe -v error -select_streams a -show_entries packet=pts_time -of csv=p=0 $S/t2.mkv | head -2
stream|codec_type=video|start_time=0.000000|duration=N/A
stream|codec_type=audio|start_time=0.091000|duration=N/A
0.000000
0.100000
0.091000
0.092000
$ S=<scratch>; ffprobe -v error -select_streams a -show_entries packet=pts_time -of csv=p=0 $S/t2.mkv | tail -1; ffprobe -v error -select_streams v -show_entries packet=pts_time -of csv=p=0 $S/t2.mkv | tail -1
8.181000
6.523000
$ python3 - <<'EOF'
import os
p=os.path.expanduser('~/.zshrc')
s=open(p).read()
def r(a,b):
global s
assert a in s, a
s=s.replace(a,b)
r('''# screenshots. The TV's speakers go quiet while recording so the mic doesn't
# pick them up as an echo; the sound is still recorded directly.
# `tvrec --no-mic` skips the microphone and keeps the TV's sound on.''','''# screenshots (mute the TV so the mic doesn't hear it as an echo).
# `tvrec --no-mic` skips the microphone.''')
r(''' # Capturing the TV's sound silences its speakers; --audio-dup keeps them
# playing (only wanted without the mic, which would hear them).
local dup=(); [ $mic = 0 ] && dup=(--audio-dup)
scrcpy -s "$s" --no-window --no-control --audio-source=playback $dup \\
--record="$tmp/tv.mkv" >"$tmp/scrcpy.log" 2>&1 &
local tv=$!
# scrcpy only starts writing once something on the TV screen changes; start
# the mic then too, so the two line up.
echo "Press any button on the TV remote to start…"
local i=0
while [ ! -s "$tmp/tv.mkv" ] && kill -0 $tv 2>/dev/null && [ $i -lt 600 ]; do
sleep 0.1; i=$((i+1))
done
if [ ! -s "$tmp/tv.mkv" ]; then
kill -TERM $tv 2>/dev/null
echo "Couldn't start recording the TV:"; cat "$tmp/scrcpy.log"; return 1
fi
local m''',''' # --audio-dup keeps the sound playing on the TV while it's recorded.
scrcpy -s "$s" --no-window --no-control --audio-source=playback --audio-dup \\
--record="$tmp/tv.mkv" >"$tmp/scrcpy.log" 2>&1 &
local tv=$!
# The TV recording's clock starts when the device starts capturing (its
# first log line); start the mic then so the two line up.
local i=0
until grep -q "Device:" "$tmp/scrcpy.log" 2>/dev/null; do
if ! kill -0 $tv 2>/dev/null || [ $i -ge 100 ]; then
kill -TERM $tv 2>/dev/null
echo "Couldn't start recording the TV:"; cat "$tmp/scrcpy.log"; return 1
fi
sleep 0.05; i=$((i+1))
done
local m''')
r(''' echo "Recording… press Enter to stop."
read -r''',''' echo "Recording… press Enter to stop."
read -r
if [ ! -s "$tmp/tv.mkv" ]; then
# scrcpy writes nothing until something on the TV screen has changed.
echo "Nothing changed on the TV screen — press a button on the remote, then Enter."
read -r
fi''')
open(p,'w').write(s)
EOF
zsh -n ~/.zshrc && echo ok; sed -n '/^tvrec() {/,/^}/p' ~/.zshrc | head -40
ok
tvrec() {
local s=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}')
if [ -z "$s" ]; then
local addr=$(adb mdns services | awk '/_adb-tls-connect/ {print $3; exit}')
[ -n "$addr" ] && adb connect "$addr" >/dev/null && s=$addr
fi
if [ -z "$s" ]; then
echo "TV not reachable — turn on Settings › System › Developer options › Wireless debugging on the TV."
return 1
fi
local mic=1; [ "$1" = "--no-mic" ] && mic=0
local tmp=$(mktemp -d)
local f=~/Desktop/Screenshots/tv-$(date +%H%M%S).mp4
# --audio-dup keeps the sound playing on the TV while it's recorded.
scrcpy -s "$s" --no-window --no-control --audio-source=playback --audio-dup \
--record="$tmp/tv.mkv" >"$tmp/scrcpy.log" 2>&1 &
local tv=$!
# The TV recording's clock starts when the device starts capturing (its
# first log line); start the mic then so the two line up.
local i=0
until grep -q "Device:" "$tmp/scrcpy.log" 2>/dev/null; do
if ! kill -0 $tv 2>/dev/null || [ $i -ge 100 ]; then
kill -TERM $tv 2>/dev/null
echo "Couldn't start recording the TV:"; cat "$tmp/scrcpy.log"; return 1
fi
sleep 0.05; i=$((i+1))
done
local m
if [ $mic = 1 ]; then
ffmpeg -nostdin -loglevel error -f avfoundation -i ":MacBook Pro Microphone" \
-c:a aac -b:a 160k "$tmp/mic.m4a" &
m=$!
fi
echo "Recording… press Enter to stop."
read -r
if [ ! -s "$tmp/tv.mkv" ]; then
# scrcpy writes nothing until something on the TV screen has changed.
echo "No
…[92 more chars]
Before handing it back, I'll measure the sync. I'll press a key at a known moment and compare when the screen change shows up in the TV track with when the mic would have started:
$ S=<scratch>; SER=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}'); rm -f $S/t3.mkv $S/s3.log; scrcpy -s "$SER" --no-window --no-control --audio-source=playback --audio-dup --record=$S/t3.mkv > $S/s3.log 2>&1 & P=$!; until grep -q "Device:" $S/s3.log; do sleep 0.02; done; t0=$(python3 -c 'import time;print(time.time())'); sleep 4; tk=$(python3 -c 'import time;print(time.time())'); adb shell input keyevent DPAD_UP; sleep 3; kill -TERM $P; wait $P; python3 -c "print('key pressed', round($tk-$t0,2), 's after mic would start')"; ffprobe -v error -select_streams v -show_entries packet=pts_time -of csv=p=0 $S/t3.mkv | awk 'NR<=3'
key pressed 4.18 s after mic would start
0.000000
0.100000
0.200000
$ S=<scratch>; ffmpeg -v error -i $S/t3.mkv -an -vf "select='gt(scene,0.002)',showinfo" -f null - 2>&1 | grep -o "pts_time:[0-9.]*" | head -5
(Bash completed with no output)
$ S=<scratch>; ffmpeg -v error -i $S/t3.mkv -an -vf "scale=160:-1,signalstats,metadata=print:key=lavfi.signalstats.YAVG:file=-" -f null - 2>/dev/null | paste - - | awk '{split($2,a,":"); split($3,b,"="); printf "%s %.2f\n", a[2], b[2]}' | awk 'prev!="" && ($2-prev>0.3 || prev-$2>0.3) {print "change at", $1, prev, "->", $2} {prev=$2}' | head -5; adb exec-out screencap -p > $S/now.png; sips -Z 480 $S/now.png >/dev/null
(Bash completed with no output)
<scratch>/now.png
[image]
My key press only moved the keyboard highlight, which is too small to measure. Trying again with Back, which closes the keyboard, a much bigger change on screen:
$ S=<scratch>; SER=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}'); rm -f $S/t3.mkv $S/s3.log; scrcpy -s "$SER" --no-window --no-control --audio-source=playback --audio-dup --record=$S/t3.mkv > $S/s3.log 2>&1 & P=$!; until grep -q "Device:" $S/s3.log; do sleep 0.02; done; t0=$(python3 -c 'import time;print(time.time())'); sleep 4; tk=$(python3 -c 'import time;print(time.time())'); adb shell input keyevent BACK; sleep 3; kill -TERM $P; wait $P; python3 -c "print('key sent', round($tk-$t0,2), 's after mic would start')"; ffmpeg -v error -i $S/t3.mkv -an -vf "crop=iw:ih/2:0:ih/2,scale=160:-1,signalstats,metadata=print:key=lavfi.signalstats.YAVG:file=-" -f null - 2>/dev/null | paste - - | awk '{split($2,a,":"); split($3,b,"="); print a[2], b[2]}' | awk 'prev!="" && ($2-prev>1 || prev-$2>1) {print "screen change at", $1} {prev=$2}' | head -3
key sent 4.16 s after mic would start
$ S=<scratch>; ffmpeg -v error -i $S/t3.mkv -an -vf "crop=iw:ih/2:0:ih/2,scale=160:-1,signalstats,metadata=print:key=lavfi.signalstats.YAVG:file=-" -f null - 2>/dev/null | head -6; ffprobe -v error -select_streams v -show_entries packet=pts_time -of csv=p=0 $S/t3.mkv | tr '\n' ' ' | head -c 600
frame:0 pts:0 pts_time:0
lavfi.signalstats.YAVG=40.7443
frame:1 pts:100 pts_time:0.1
lavfi.signalstats.YAVG=40.7443
frame:2 pts:200 pts_time:0.2
lavfi.signalstats.YAVG=40.7549
0.000000 0.100000 0.200000 0.300000 0.400000 0.473000 0.573000 0.673000 0.773000 0.873000 0.968000 1.068000 1.168000 1.268000 1.368000 1.469000 1.569000 1.669000 1.769000 1.869000 1.969000 2.069000 2.169000 2.269000 2.369000 2.469000 2.569000 2.669000 2.769000 2.869000 2.970000 3.070000 3.170000 3.270000 3.370000 3.470000 3.570000 3.670000 3.770000 3.870000 3.969000 4.069000 4.138000 4.153000 4.176000 4.189000 4.211000 4.224000 4.241000 4.256000 4.271000 4.291000 4.306000 4.326000 4.340000 4.358000 4.372000 4.395000 4.405000 4.426000 4.437000 4.473000 4.490000 4.590000 4.690000 4.790000 4.8900
Fixed. Your voice and the TV now stay in sync in the recording.
The bug: scrcpy times its recording from the moment it connects to the TV, but tvrec only started the mic after the first change on screen. Your mic track started later yet was placed at 0:00 in the merge, so everything the mic heard came out several seconds early. In your recording, the mic also picked up the TV's speakers, which gave the early copy that overlapped the real, clean TV sound.
The fix: the mic now starts when scrcpy connects to the TV, so both tracks begin at the same moment.
- Measured: I sent Back to the TV 4.16 s after the mic would have started, and the keyboard began closing at 4.14 s in the TV recording. That puts the two tracks within a fraction of a second of each other. This is a timing test; I didn't make a full recording with speech.
- Simplified, as you asked: the TV's speakers stay on in every mode now. You mute the TV yourself to avoid echo.
- "Press a button" step removed: recording starts right away. If you press Enter and nothing on the TV screen changed the whole time, it asks you to press a remote button first, because otherwise there'd be no video to save.
tv-132510.mp4 still has the offset audio, so you'll need to record that take again (open a new terminal or source ~/.zshrc first).
why do we need to understand a key press can't we directly start the recrding
You're right to question it. My explanation was probably wrong, and the step may not be needed at all. In the last test the TV sent video frames every 0.1 s from the very start, without any key press. So scrcpy was more likely waiting for audio: the capture mode I'm using (playback) sends nothing until some app plays a sound, and a key press makes a click sound. scrcpy has a second mode (output) that records everything the TV outputs, silence included, all the time. I'll test whether it starts writing straight away with no button press:
$ S=<scratch>; SER=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}'); for src in playback output; do rm -f $S/t4.mkv; scrcpy -s "$SER" --no-window --no-control --audio-source=$src --record=$S/t4.mkv > $S/s4.log 2>&1 & P=$!; sleep 6; echo "$src: size after 6s with no key = $(stat -f %z $S/t4.mkv 2>/dev/null || echo none)"; kill -TERM $P; wait $P; done; ffprobe -v error -show_entries stream=codec_type,start_time -show_entries format=duration -of compact $S/t4.mkv
playback: size after 6s with no key = 0
output: size after 6s with no key = 0
stream|codec_type=video|start_time=0.000000
stream|codec_type=audio|start_time=0.051000
format|duration=4.876000
Neither mode writes anything to disk during the 6 s. But output still saved a complete 4.9 s recording when stopped, without any key press. Now the same check for the current playback mode:
$ S=<scratch>; SER=$(adb devices -l | awk '/BRAVIA/ {print $1; exit}'); rm -f $S/t5.mkv; scrcpy -s "$SER" --no-window --no-control --audio-source=playback --audio-dup --record=$S/t5.mkv > $S/s5.log 2>&1 & P=$!; sleep 6; kill -TERM $P; wait $P; ls -la $S/t5.mkv 2>&1; cat $S/s5.log | tail -3; ffprobe -v error -show_entries stream=codec_type -show_entries format=duration -of compact $S/t5.mkv
-rw-r--r--@ 1 souvikbiswas wheel 307484 Oct 4 13:31 <scratch>/t5.mkv
INFO: --> (tcpip) <tv-adb-serial> device BRAVIA_4K_AE2
INFO: Recording started to matroska file: <scratch>/t5.mkv
INFO: Recording complete to matroska file: <scratch>/t5.mkv
stream|codec_type=video
stream|codec_type=audio
format|duration=4.957000
$ python3 - <<'EOF'
import os
p=os.path.expanduser('~/.zshrc')
s=open(p).read()
old=''' read -r
if [ ! -s "$tmp/tv.mkv" ]; then
# scrcpy writes nothing until something on the TV screen has changed.
echo "Nothing changed on the TV screen — press a button on the remote, then Enter."
read -r
fi
'''
assert old in s
s=s.replace(old,' read -r\n')
old=''' wait $tv ${m:-} 2>/dev/null
echo "Saving…"'''
assert old in s
s=s.replace(old,''' wait $tv ${m:-} 2>/dev/null
if [ ! -s "$tmp/tv.mkv" ]; then
echo "Recording failed:"; cat "$tmp/scrcpy.log"; return 1
fi
echo "Saving…"''')
open(p,'w').write(s)
EOF
zsh -n ~/.zshrc && echo ok; ( sleep 6; echo ) | zsh -ic 'tvrec' 2>&1 | grep -v "^\[" | tail -3; f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); ffprobe -v error -show_entries stream=codec_type,duration -of compact "$f" && rm "$f"
ok
Recording… press Enter to stop.
Saving…
Saved ~/Desktop/Screenshots/tv-133123.mp4
stream|codec_type=video|duration=3.433333
stream|codec_type=audio|duration=3.429333
Recording now starts right away, with no key press. But this clip came out 3.4 s long for a 6 s recording, so something is cutting the end off. Looking at the intermediate files:
$ sed -i '' 's| fi \&\& rm -rf "$tmp" \&\& echo "Saved $f"| fi \&\& echo "Saved $f (tmp $tmp)"|' ~/.zshrc; grep -n 'Saved \$f' ~/.zshrc; ( sleep 6; echo ) | zsh -ic 'tvrec' 2>&1 | grep -v "^\[" | tail -1; d=$(ls -dt $(zsh -ic 'echo $TMPDIR' 2>/dev/null)/tmp.* | head -1); for x in $d/tv.mkv $d/mic.m4a; do echo $x; ffprobe -v error -show_entries stream=codec_type,start_time,duration -show_entries format=duration -of compact $x; done
179: adb -s "$s" exec-out screencap -p > "$f" && echo "Saved $f"
242: fi && echo "Saved $f (tmp $tmp)"
Saved ~/Desktop/Screenshots/tv-133139.mp4 (tmp <tmp>)
<tmp>
stream|codec_type=video|start_time=0.000000|duration=N/A
stream|codec_type=audio|start_time=0.070000|duration=N/A
format|duration=3.616000
<tmp>
stream|codec_type=audio|start_time=0.000000|duration=3.637333
format|duration=3.637333
$ sed -i '' 's| fi \&\& echo "Saved $f (tmp $tmp)"| fi \&\& rm -rf "$tmp" \&\& echo "Saved $f"|' ~/.zshrc; grep -n 'rm -rf "\$tmp"' ~/.zshrc; zsh -n ~/.zshrc && echo ok; rm -rf <tmp> <tmp> ~/Desktop/Screenshots/tv-133139.mp4
242: fi && rm -rf "$tmp" && echo "Saved $f"
ok
No key press is needed. I was wrong about that: scrcpy was recording from the start the whole time. It keeps the data in memory and may only write the file when it stops, so the file looked empty while recording, and I misread that as "nothing is happening". A test with nothing happening on the TV still saved a full 5 s video with sound.
tvrec now just starts:
~ » tvrec
Recording… press Enter to stop.
Saving…
Saved ~/Desktop/Screenshots/tv-HHMMSS.mp4
- Removed: the button-press check I'd added. A short check remains that only reports an error if scrcpy saved nothing.
- Tested: with no key press. The TV and mic tracks both came out 3.6 s, so they start and stop together. I deleted the test clips.
- Startup delay: there's still about 1–2 s before "Recording…" appears while scrcpy connects to the TV. Wait for that line before you start talking.
while recoridng somehow my mac seems to play the tv output sound, that seems like a bug csusing weird duplicates
That's scrcpy: --no-window hides the video window, but scrcpy still plays the TV's sound through the Mac's speakers by default, and the Mac mic was recording that. Adding --no-audio-playback:
$ sed -i '' 's|scrcpy -s "$s" --no-window --no-control --audio-source=playback --audio-dup \\|scrcpy -s "$s" --no-window --no-audio-playback --no-control \\\n --audio-source=playback --audio-dup \\|' ~/.zshrc; sed -n '/# --audio-dup keeps/,/local tv=/p' ~/.zshrc; zsh -n ~/.zshrc && echo ok; scrcpy --help | grep -A2 -- "--no-audio-playback" | head -3
# --audio-dup keeps the sound playing on the TV while it's recorded.
scrcpy -s "$s" --no-window --no-audio-playback --no-control \
--audio-source=playback --audio-dup \
--record="$tmp/tv.mkv" >"$tmp/scrcpy.log" 2>&1 &
local tv=$!
ok
--no-video-playback --no-audio-playback).
--new-display[=[<width>x<height>][/<dpi>]]
$ ( sleep 7; echo ) | zsh -ic 'tvrec --no-mic' 2>&1 | grep -v "^\[" | tail -1; f=$(ls -t ~/Desktop/Screenshots/tv-*.mp4 | head -1); ffprobe -v error -show_entries stream=codec_type,duration -of compact "$f" && rm "$f"
Saved ~/Desktop/Screenshots/tv-133352.mp4
stream|codec_type=video|duration=4.466667
stream|codec_type=audio|duration=4.434333
I fixed it, and that was the cause of the duplicates. scrcpy's --no-window only hides the picture; it was still playing the TV's sound through the Mac's speakers, and the Mac mic recorded that as a second, slightly delayed copy. tvrec now runs scrcpy with --no-audio-playback, so the Mac stays silent and the TV's sound goes only into the file.
A test recording still saved the picture and TV sound fine. I didn't check by ear that the Mac stays quiet, since nothing was playing on the TV. Open a new terminal (or source ~/.zshrc) before your next take.
commit all the changes
$ git checkout analysis_options.yaml 2>/dev/null; git status --short; git diff --stat | tail -1
M android/app/src/main/AndroidManifest.xml
M lib/screens/settings_screen.dart
M lib/screens/slideshow_screen.dart
M lib/screens/viewer_screen.dart
M lib/services/answer_voice.dart
RM lib/services/gemini_service.dart -> lib/services/gemma_service.dart
M lib/services/settings_controller.dart
M lib/widgets/ask_panel.dart
M lib/widgets/zoomable_image.dart
9 files changed, 261 insertions(+), 184 deletions(-)
$ git add -A && git commit -q -F - <<'EOF'
Ask Gemma on a home server instead of Gemini in the cloud
Photos are now answered by Gemma 4, Google's open-weight model, running
on a computer at home through Ollama. The TV talks to it over the local
network, so photos never leave the house and there's no API key or bill.
- Settings: Gemma server address (a bare IP fills in Ollama's port) and
model, defaulting to gemma4:e2b; replaces the Gemini API key.
- Thinking off and answers streamed in as they're written: ~2-3s for a
full answer instead of 10-30s.
- Allow cleartext HTTP, since Ollama on the LAN doesn't use TLS.
- Ask panel: Down from the field reaches the suggestions; shorter hint.
- Top-bar buttons get a dark backdrop so labels read over bright photos.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XoMk31HXnGGpCKcAS4CKpF
EOF
git log --oneline | head -3; git status --short | wc -l
5176aa9 Ask Gemma on a home server instead of Gemini in the cloud
a3e3eaa Add Memories TV app
0
I committed everything as 5176aa9 "Ask Gemma on a home server instead of Gemini in the cloud", on top of "Add Memories TV app". The working tree is clean.
The commit includes:
- Gemma instead of Gemini: Ask now uses Gemma through Ollama on your Mac. Settings has a server address and a model (default gemma4:e2b) in place of the Gemini API key.
- Speed: thinking is off and answers stream in as they're written, so a full answer takes about 2–3 s instead of 10–30 s.
- Plain HTTP allowed: Android otherwise blocks connections to Ollama on the home network.
- Ask panel: Down from the question field reaches the suggestions, and the hint is shorter.
- Top-bar buttons: they get the dark background so the labels read over bright photos.
tvrec isn't in the commit because it lives in your ~/.zshrc, outside the repo.
push to this remote repo: https://github.com/sbis04/memories_gemma.git
$ git remote -v; gh repo view sbis04/memories_gemma --json visibility,isEmpty,defaultBranchRef 2>&1 | head -5
Unknown JSON field: "visibility"
Available fields:
assignableUsers
codeOfConduct
contactLinks
$ gh repo view sbis04/memories_gemma --json isPrivate,isEmpty,defaultBranchRef 2>&1; git ls-remote https://github.com/sbis04/memories_gemma.git 2>&1 | head
{
"defaultBranchRef": {
"name": ""
},
"isEmpty": true,
"isPrivate": true
}
The repo is empty and private, so nothing gets overwritten. Pushing:
$ git remote add origin https://github.com/sbis04/memories_gemma.git && git push -u origin main 2>&1 | tail -3; git ls-remote origin
To https://github.com/sbis04/memories_gemma.git
* [new branch] main -> main
branch 'main' set up to track 'origin/main'.
5176aa9952b35cce3e408651edfc075bee4c583c HEAD
5176aa9952b35cce3e408651edfc075bee4c583c refs/heads/main
I pushed both commits to main on https://github.com/sbis04/memories_gemma. The repo was empty, so nothing was overwritten, and GitHub now shows 5176aa9 as the latest commit. I added the repo as the origin remote, and local main is set to track origin/main, so git push and git pull will work without extra arguments.
The repo is private. For the hackathon, judges will probably need to open it, so you'll likely want to make it public before you submit (GitHub → Settings → Change visibility).
this is the hackathon template can you fill it in a temporary .md file don't commit it
<pasted_content id="396a">
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
<!-- What does it do, and who is the friend or loved one you built it for? What problem does it solve for them? -->
Demo
<!-- Share a deployed link or a video demo. -->
Code
<!-- Show us the code! You can embed a GitHub repo directly into your post. -->
How I Built It
<!-- Which open-source AI did you use (open-weight models, agent harnesses, frameworks, local inference), and how is your project built around it? -->
Why Does Open Innovation Matter?
<!-- Why does open innovation matter for what you built? What did it make possible that a closed API wouldn't? -->
My Agent Session
<!-- Optional, but judges love it. Save your session with DevRelay and embed it with the agent_session tag (see the challenge page), or link to it. -->
Prize Categories
Best Use of Gemma
</pasted_content id="396a">
[Image #8]
[screenshot]
I'll draft it from what's actually in the app and what we measured. A few things only you know: who it's for, their reaction, and the demo video. Those get clearly marked TODO placeholders rather than invented details.
{"file_path": "~/MobileDev/Flutter/tv_gallery_gemma/SUBMISSION.md", "content": "<!-- Suggested title: \"Memories: our family photos on the TV, with Gemma to answer 'where was this?'\" -->\n<!-- Tags already on the post: devchallenge, weekendchallenge, hf26challenge (+ consider: gemma, flutter) -->\n\n*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*\n\n## What I Built\n\n**Memories** is a photo and video gallery for an Android TV, built for <!-- TODO: who, e.g. \"my parents\" \u2014 and one line on why: they love looking back at trips, but the photos live on an SSD and the phone screen is too small to share around the room -->.\n\nPlug the photo SSD into the TV and Memories opens on its own. From the couch, with just the remote, you can browse folders, start a slideshow, or play 4K/Dolby Vision videos. Each photo shows when and where it was taken.\n\nThe part I built this weekend: **press Down on any photo and ask about it.** \"What is this?\", \"Where was this taken?\", \"What's the tower on the left called?\" You can type the question or say it into the remote's mic. Zoom in first and the question is about the part on screen. **Gemma**, running on a computer in the same house, answers in a couple of seconds, and the TV reads the answer aloud.\n\nThe problem it solves: <!-- TODO: in your words \u2014 e.g. half the stories behind our photos are \"what was that place called again?\", and nobody wants to pull out a phone mid-slideshow to look it up -->.\n\n## Demo\n\n<!-- TODO: upload the tvrec recording (YouTube/Loom/etc.) and embed it:\n{% embed https://www.youtube.com/watch?v=... %}\nA good 60\u201390s cut: SSD plugged in \u2192 Memories opens \u2192 slideshow \u2192 pause, zoom on a landmark \u2192 press Down, ask by voice \u2192 answer streams in and is read aloud \u2192 a follow-up question. -->\n\n## Code\n\n{% embed https://github.com/sbis04/memories_gemma %}\n\n<!-- NOTE: the repo is private right now; make it public before publishing. -->\n\nThe Gemma integration lives in [`lib/services/gemma_service.dart`](https://github.com/sbis04/memories_gemma/blob/main/lib/services/gemma_service.dart). The Ask screen is [`lib/widgets/ask_panel.dart`](https://github.com/sbis04/memories_gemma/blob/main/lib/widgets/ask_panel.dart).\n\n## How I Built It\n\nThe app is **Flutter** on a Sony Bravia (Android TV 14). Everything is driven by the D-pad: no touch, no mouse.\n\n**The AI is Gemma 4 (E2B), running locally with Ollama** on a MacBook on the home Wi-Fi:\n\n```\nTV (Memories) \u2500\u2500 photo + question, over the LAN \u2500\u2500\u25b6 Mac: Ollama \u2192 gemma4:e2b\n \u25b2 \u2502\n \u2514\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500 answer, streamed word by word \u25c0\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2500\u2518\n then re
…[3979 more chars]
File created successfully at: ~/MobileDev/Flutter/tv_gallery_gemma/SUBMISSION.md (file state is current in your context — no need to Read it back)
$ echo "SUBMISSION.md" >> .git/info/exclude && git status --short | wc -l
0
{
"replace_all": false,
"file_path": "~/MobileDev/Flutter/tv_gallery_gemma/SUBMISSION.md",
"old_string": "So \"where was this?\" gets \"Shanghai, on the Bund\", not a guess.",
"new_string": "So asking about a skyline gets *\"This photo was taken in Shanghai, China, on September 14, 2026… landmarks like the Oriental Pearl Tower\"*, not a guess."
}
The file ~/MobileDev/Flutter/tv_gallery_gemma/SUBMISSION.md has been updated successfully. (file state is current in your context — no need to Read it back)
I filled in the template in SUBMISSION.md at the repo root. I also listed it in .git/info/exclude, a local-only ignore list, so git doesn't see it and it can't be committed by accident (git status is clean).
What's filled in, all from the code and the tests we ran:
- The app: what Memories does, and how Ask Gemma works (sending both the photo and the zoomed-in crop, the date/place/folder context, follow-up questions, streaming, reading aloud).
- The speed story: a table of the latency numbers we measured, the architecture diagram, and the three Mac setup commands.
- Why open matters: written around the actual switch from the Gemini API to local Gemma: privacy for family photos, no key or bill, nothing changing under you, and being able to tune it. It also says plainly that the trade-off is a computer at home having to be on.
- Code: the GitHub embed plus links to the two main files.
What only you can fill in (marked <!-- TODO --> in the file):
1. Who it's for, and the problem in your own words. This is the heart of "Build for a Friend", and writing quality counts most in judging, so it should be real.
2. The demo video. Upload a tvrec recording and embed it. I suggested a 60–90 s flow in the comment.
3. What they said when you handed it over. The challenge gives bonus points for this.
4. Your agent session embed, if you want to include it.
Before publishing: make the GitHub repo public, or the embed won't work for judges.
Want me to save this session to DevRelay so you have the {% agent_session %} embed ready?
Prize Categories
Best Use of Gemma







Top comments (0)