DEV Community

Souvik Biswas
Souvik Biswas

Posted on

Memories: Personal Media Viewer for TV with local visual intelligence

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

Memories is a photo and video viewer for our Android TV. I built it for especially my mom who loves to browse through old photos and videos on TV 🙂

Memories start screen

Just last week my mom asked if there's an app that can analyze a picture and tell me where it was taken. I thought that would be a fun weekend project to try out Gemma for visual intelligence. Now, this app has that capability plus a lot more!

You plug the photo SSD into the TV and the app opens by itself. Everything works from the couch with the regular remote. You can browse folders and run slideshows. It even plays the 4K Dolby Vision videos from our trips. Every photo also shows the date and the place it was taken. One of the other reasons why I started working on this app was the sheer lack of a good photo viewer for Android TV. The built-in one is slow and clunky. I wanted something that was fast, simple, and fun to use.

Photo & Video view

This weekend I added the part I'm most excited about. On any photo you can press Down on the remote and just ask about it. Things like "What is this place called?" or "What's that tower on the left called?". You can type the question or hold the mic button and say it.

Gemma View

If you zoom into part of the photo first then the question is about that part. The answer comes from Gemma running on a Mac in the same house. It shows up in a couple of seconds and the TV reads it out loud.

Demo

Code

GitHub logo sbis04 / memories_gemma

Media viewer for Android TV

Memories

A photo and video viewer for Android TV that lets you ask questions about your photos. The answers come from Gemma running on a computer in your home, so your photos never leave the house.

Memories start screen

Features

  • Made for the remote. Plug in your photo drive and the app opens by itself. Browse folders as a grid or list and pin the ones you use most.
  • Slideshows with fade, Ken Burns or slide transitions. They resume where you left off.
  • 4K, HDR and Dolby Vision video through the TV's own decoder.
  • Date and place on every photo, read from its GPS data.
  • Ask Gemma. Press Down on a photo and ask about it by typing or with the remote's mic. Zoom in first to ask about one part. The answer streams in and the TV reads it aloud.
Memories.Gemma.Demo.mov
A folder of trip videos in the gallery grid Folder menu with Pin to top, Rename and Delete
Photos and videos laid out to fill the screen Pin, rename
…

How I Built It

The app is written in Flutter and runs on a Sony Bravia with Android TV 14. There's no touch screen and no mouse. Every screen had to work with just the D-pad.

For the AI part I'm using Gemma 4 (the E2B size) through Ollama. It runs on my MacBook and the TV talks to it over the home Wi-Fi:

TV app  ── photo + question over LAN ──▶  Mac: Ollama → gemma4              ▲                                                             │
└──────── answer, streamed word by word ◀───────────────┘
   then read aloud by the TV's built-in text-to-speech
Enter fullscreen mode Exit fullscreen mode

Settings for Gemma Server

Here's roughly what happens when you ask something:

  • The app sends Gemma the photo shrunk to about 1024px. If you're zoomed in it also sends a sharper crop of what's on screen. Gemma is told to focus on that crop. That's what makes "what's that building?" work on a wide skyline shot.
  • The first question also carries what the app already knows about the photo. That's the date and the place name from the photo's GPS and the album folder. So for our Shanghai photo Gemma came back with "This photo was taken in Shanghai, China, on September 14, 2026… landmarks like the Oriental Pearl Tower" instead of guessing.
  • Follow-up questions keep the earlier ones as context. So "and the one next to it?" works.
  • The answer streams in over Ollama's /api/chat as Gemma writes it. Once it's done the TV reads it out with Android's built-in voice.

Speed took the most work. My first version took anywhere from 10 to 30 seconds per answer. On a TV that feels like the app has frozen. When I looked at where the time went I found Gemma 4 spending around 500 tokens "thinking" before answering even "what is this?". I turned thinking off (think: false) and switched to the smaller E2B model. For photo questions I honestly couldn't tell the difference in the answers:

Setup (M3 Max, model already loaded) Full answer First words on screen
gemma4 (E4B), thinking on 10 to 29 s nothing until the whole answer is done
gemma4 (E4B), thinking off about 5 s under 1 s (streamed)
gemma4:e2b, thinking off about 2.5 to 3 s under 1 s

I also ask Ollama to keep the model in memory for 30 minutes. That way only the first question of the evening has to wait for it to load.

If you want to try it then the Mac side is just this:

brew install ollama
ollama pull gemma4:e2b
OLLAMA_HOST=0.0.0.0 ollama serve   # so the TV can reach it over Wi-Fi
Enter fullscreen mode Exit fullscreen mode

Then put the Mac's IP address in Memories under Settings > Gemma server.

Why Does Open Innovation Matter?

The first version of this feature actually used a cloud API. Every question sent the photo off to a hosted model. You also needed an API key to use it. Switching to Gemma changed a few things that matter a lot for an app like this.

The big one is privacy. These are our family photos. They have kids and faces in them along with our house and every place we've lived. Now when you ask about a photo it only travels from the TV to the Mac in the next room. Nothing gets uploaded and I don't have to read anyone's data policy.

It's also free to run. Nobody has to type an API key with a TV remote. There's no bill or rate limit to think about either.My parents will never need to know what a token is 😂

It won't change on me. The cloud model I started with sat behind a "latest" alias. That's nice right up until it starts answering differently one day. The Gemma weights are sitting on my disk. If it works today it should keep working.

Running it locally also meant I could dig into it. I found the thinking problem myself and measured every change I made. Picking the smaller model and adding streaming came out of that. Trying another model is just a setting in the app. I use gemma4:e2b for speed. gemma4 gives slightly nicer wording and gemma4:26b is there if you have a beefier machine.

The catch is that a computer at home has to be switched on. We mostly use the TV in the evening and the Mac is right there anyway. I'm happy with that if it means our photos stay at home.

My Agent Session

I did the Gemma work in one afternoon with Claude Code. I swapped the cloud API for Ollama and set up Gemma on the Mac. Then I chased down the slow answers and tested everything on the actual TV. I even had it write the little script I used to record the demo.

Memories: Personal Media Viewer for TV with Gemma for visual intelligence
You

can we add gemma to this project to make it eligible for this hackathon: https://dev.to/challenges/hacktoberfest-weekend-2026-10-01

[Image #1] [Image #2] [Image #3] [Image #4] [Image #5]
[screenshot]
[screenshot]
[screenshot]
[screenshot]
[screenshot]

Agent

Going with Gemma on a home Mac through Ollama, as you picked. I'm now swapping the app's Gemini code over to call it.

You

actually first commti the inital app before making these new cahnges, create a commit called "Add Memories TV app"

Agent

The git history has been reset (no commits yet). First I'll undo my file rename so the commit has the original app, then check that nothing secret or generated gets staged.

Now the settings: replacing the API key with a server address and a model name.

"Ask Gemini" is now "Ask Gemma". Gemma runs on a computer at home through Ollama, and the TV sends it the photo over the home Wi-Fi. flutter analyze passes, but I haven't sent it a real question yet because Ollama isn't installed on your Mac. These changes aren't committed; the commit before them is "Add Memories TV app".

What changed
- lib/services/gemma_service.dart (renamed from gemini_service.dart): it now calls Ollama's /api/chat instead of Google's API. It still sends the downscaled photo, the zoomed-in crop, and the date, place and folder, and follow-up questions still work.
- It asks Ollama to keep the model loaded for 30 minutes, so later questions don't wait for it to reload.
- The first question can take up to 3 minutes while the model loads.
- Error messages say what to fix: Mac off or Ollama not on the network, model not downloaded (with the exact ollama pull command), or the model being too slow.
- Settings: the Gemini API key is replaced by "Gemma server" and "Model".
- For the server you can type just an IP or name like [REDACTED] and it fills in http://…:11434.
- The model defaults to gemma4, Gemma 4's smaller default size (~9.6 GB).
- Text on screen: every "Gemini" is now "Gemma". The "no server set" screen explains the setup in one sentence.
- AndroidManifest.xml: I turned on plain-HTTP traffic. Android blocks it by default, and Ollama doesn't use HTTPS on a home network, so without this the TV couldn't connect.

Setting up the Mac
sh
brew install ollama
ollama pull gemma4 # or gemma4:26b — fits in your 36 GB and answers better
OLLAMA_HOST=[REDACTED] ollama serve # listen on the LAN, not just localhost

Then put the Mac's IP into the TV's Settings.

For your post: the photos stay in the house, there's no API key or bill, and you can switch models by changing one setting. That covers both "keep data off a server they don't control" and "cost nothing to run". The one catch is that the Mac has to be on while you use it.

Do you want me to install Ollama and download gemma4 (about 10 GB) so I can test a real question end to end before we commit?

Sources:
- How to Run Gemma 4 with Ollama (codersera)
- Best Ollama Models 2026 (morphllm)

You

can you do that mac setup

Agent

Ollama is running as a background service. Next I'm making it listen on the home network so the TV can reach it (by default it only accepts connections from the Mac itself):

Ollama is now listening on all interfaces. Downloading gemma4 (~10 GB) next:

Model downloaded. Now I'll send it the same request the TV sends, with a photo, through the Mac's network address instead of localhost:

The test went through before your reply arrived. Since you've accepted the dialog, I'll run it again to check the first-answer time:

Your Mac is set up, and Gemma answers questions about a photo over the home network.

What I did
- Installed Ollama 0.35.1 with Homebrew. It runs as a background service that starts when you log in.
- Set OLLAMA_HOST=[REDACTED] in the service's launch file (~/Library/LaunchAgents/sh.brew.ollama.plist), so it accepts connections from other devices on the Wi-Fi.
- Downloaded gemma4 (6.6 GB) and checked that it runs on the M3 Max's GPU.
- The dialog you accepted was the macOS firewall. It now shows incoming connections to Ollama as permitted.

Test results (I sent the app's exact request to http://[REDACTED]:11434 with a 1024px JPEG):
- First question: took 149 s. That was the model loading for the first time, and probably the firewall dialog waiting on you too.
- Same question again, model already loaded: 10.6 s.
- Follow-up question: 9.5 s, and the answer used the earlier conversation.
- Answers: accurate, plain text, 2–5 sentences as the app asks for, and they used the folder and date I passed in.
- Model name that doesn't exist: returns 404 with model '…' not found, which the app turns into the "run ollama pull …" message.

These requests came from the Mac to its own network address. The TV itself hasn't connected yet.

On the TV: go to Settings → Gemma server and enter [REDACTED]. The app fills in the port.

Two things to watch
- The Mac's IP can change. If the router hands out a new address, the TV can't find Gemma. Set a fixed (reserved) address for the Mac in your router, or use its .local name.
- brew services restart ollama undoes the network setting. It rewrites the launch file without OLLAMA_HOST, so after it Ollama only listens to the Mac itself. Plain reboots are fine. If you run that command, tell me and I'll put the setting back.

The Gemma code changes are still not committed. Should I commit them now, or first build the app onto the TV and test it there?

You

Can we use faster model to make it reply under 5 seconds or even less

Agent

I'll first measure where the 10 s goes (reading the image, thinking, or writing the answer), then compare models and settings.

Code updated and analyzer clean. Now I'm timing the streaming request against the real Gemma server: when the first words arrive and when the answer finishes.

Answers are now about 2–3 s from asking to the full answer on screen with the default gemma4:e2b, and the first words show up in under a second. All of this was measured on the Mac against its own network address. I haven't run it on the TV yet.

What I changed
- Thinking off: this was the main cause. Gemma 4 was silently reasoning for about 500 words before every answer, which took 10–29 s. Turning it off (think: false) costs almost nothing in quality for describing a photo.
- Default model is now gemma4:e2b: I downloaded it next to gemma4, which stays one setting away under Settings → Model if you prefer slightly better wording.
- Answers stream in: the words appear on screen as Gemma writes them, instead of all at once at the end. If something fails halfway through, the partial answer is replaced by the error message. Reading aloud still starts once the full answer is in.

Measured (photo already prepared, model loaded)

Full answer First words
gemma4:e2b ~2.5–3 s ~0.1–0.9 s
gemma4 ~5 s ~0.1–0.9 s*
Before (thinking on) 10–29 s same as full answer

*The first gemma4 question after using e2b took about 6 s before any words appeared, because the Mac had to load the model. That only happens when you switch models or after 30 minutes idle.

In the streaming test I left out the app's short-answer instruction, so the answers ran long and slow there (4–13 s). With the instruction, as in the earlier test, they're 2–5 sentences, which gives the times in the table.

flutter analyze passes. Should I commit now, or build onto the TV first so you can try it there?

You

build and install on the TV

Agent
You

let's do debug for now

Agent

No TV is connected over adb right now. I'll check whether it shows up on the network:

I can't find the TV on the network. None of the devices answers on adb's port (5555), and nothing advertises wireless debugging. I need you to turn on debugging on the Bravia:

  1. Go to Settings → System → About and click Build seven times to unlock developer options (skip this if you've done it before).
  2. Go to Settings → System → Developer options and turn on USB debugging, plus Network debugging / ADB over network if it's there.
  3. Send me the TV's IP address from Settings → Network & Internet → (your Wi-Fi) → IP address.

Once you send the IP, I'll run adb connect <ip>. You'll need to accept an "Allow USB debugging?" prompt on the TV, then I'll start the debug build with flutter run. If you'd rather connect it yourself, type ! adb connect <tv-ip> here and I'll take it from there.

You

chedck now

Agent

The TV is connected (BRAVIA 4K AE2 at [REDACTED]). Starting the debug build on it:

The debug build is running on the TV, the BRAVIA 4K AE2, connected over Wi-Fi. I'm watching its log for errors while you try it.

To try it:
1. In Memories, open Settings → Gemma, go to Gemma server → Add, and enter [REDACTED] (the Mac). Leave Model as gemma4:e2b.
2. Open a photo and press Down on the remote to open Ask Gemma, then pick a suggestion or ask something.

The first question can take about 5 s while the Mac loads the model. After that, answers should start appearing within about a second.

Tell me how it goes. I can hot-reload any changes straight onto the TV.

You

'~/Desktop/Screenshots/tv-125724.png'

remove the "-- or use the keyboard mic" text part in hint

Agent

The hint now just says "Ask anything about this photo". I hot-reloaded it onto the TV, so it shows next time Ask Gemma opens.

You

seems like even with the keyboard closed, I can't hit down to selct any of the options from there

Agent

Down is being swallowed by the text field: it handles Down itself, as "move the cursor to the end". Once there's a conversation, Down already moves into the answers. Before any question, it should go to the suggestion pills instead:

I fixed it and hot-reloaded it onto the TV. With the keyboard closed and nothing asked yet, Down from the question field now jumps to the suggestion pills, and Up should take you back to the field. After you've asked something, Down still moves into the answers to scroll them. Close and reopen Ask Gemma to try it.

You

closing the keyboard with back shoudl actually unfocus the text field as well I think

Agent

First I'll check how the TV reports the keyboard closing. I can drive the TV over adb and take screenshots:

I'm checking how the TV signals the keyboard closing. A debug probe I added caused an error (I didn't unregister it when the panel closes), so I'm looking at that first:

Okay, I'll leave it. Removing my debug probe:

I'm leaving the Back/keyboard behaviour as it is. I removed the temporary debug code (it was what caused the error in the log) and hot-reloaded the TV. I also reverted the analysis_options.yaml change that flutter run added on its own.

If you change your mind later, here's what I found: the TV keyboard never tells the app it's open (the space it covers always reads 0). Android does still mark it as visible, so the app could ask Android directly through a small piece of Android code.

Nothing is committed yet. The changes are the switch to Gemma, the speed changes, the shorter hint, and Down reaching the suggestion pills. The debug app is still running on the TV. Say when you want me to commit.

The old panels are still registered for my removed debug hook in the running app, so they keep throwing. A hot restart clears them:

Those alerts are from before the restart arriving late. To confirm:

The errors have stopped. They came from my temporary debug code still being hooked into panels that were already open, and a hot restart cleared them. The log has been clean since the restart, and the TV app reopened where you were.

The rest is unchanged: nothing is committed yet, and the debug app is still running on the TV.

You

'~/Desktop/Screenshots/tv-130502.png'for the top bar buttons can you add a 50% opacity dash background color to make sure the texts are always reasdable

Agent

I'm taking "dash" to mean a dark background. Finding the top-bar chip widget:

The controls have auto-hidden. Pressing Up to bring them back:

I gave the top-bar buttons a 50% black background when they're not selected. The selected button stays solid white. I assumed "dash" was a typo for "dark". It's hot-reloaded on the TV. In the screenshot I took, "Ask" now reads clearly over the bright sky, where before it nearly disappeared.

I changed the same thing on the slideshow's top bar (the Play/Pause and Ask buttons) so the two screens match. Those were faint white at 14% before.

You

is ther a way I can record the screen of the TV? from this mac

I already have an alis to capture screenshot using tvshot

Agent

Yes. Android TV has a built-in screen recorder (adb shell screenrecord), so you don't need to install anything. I'll add a tvrec command next to your tvshot, in the same style:

A test recording saved. I'm checking that the file is a complete, playable video:

The file is 1080p but only 0 s long. The TV screen didn't change during the test, and Android's recorder only saves frames when something on screen moves. I'll test again while moving through photos:

Yes. I added a tvrec command to your ~/.zshrc, right below tvshot, and it works.

tvrec        # starts recording; press Enter to stop

It finds the TV the same way tvshot does and saves ~/Desktop/Screenshots/tv-HHMMSS.mp4 at 1080p. It uses the recorder built into Android TV, so nothing extra is installed. Open a new terminal, or run source ~/.zshrc, to get it.

  • Tested: I recorded while pressing the remote buttons over adb and got a 6.8 s, 1920×1080 video. The test clip is at ~/Desktop/Screenshots/tv-130753.mp4 if you want to look at it or delete it.
  • Limits: a single recording stops at 3 minutes, and there's no sound. A screen that doesn't change records as almost nothing; my first test on a still photo came out 0 s long.
  • If you need sound or longer clips (for example the voice reading answers aloud in your demo): brew install scrcpy, then scrcpy --record demo.mp4. It mirrors the TV in a window on the Mac and records audio too, on Android 11 and later. I haven't installed it. Tell me if you want it.
You

can it record voice that I speak during the input and the output voice too?

Agent

The TV's own sound (Gemma's answers read aloud) can be recorded. Your voice is harder, because it goes into the remote's mic and only the TV's voice search hears it.

tvrec (now) scrcpy Mac microphone
Video ✅ ✅ —
TV sound / spoken answers ❌ ✅ likely ✅ if the Mac is in the room
Your voice ❌ ⚠️ probably not ✅
  • TV sound: scrcpy can grab everything the TV plays, which would include the spoken answers. I haven't tested it with this TV's text-to-speech yet. That's the first thing to check.
  • Your voice through scrcpy: it can record from the TV's microphone (--audio-source=mic). But while you hold the mic button, Google's voice typing is using that mic, and Android usually gives other recorders silence then. It's also unclear whether a Bluetooth remote's mic counts as the TV's mic at all. I expect this not to work.
  • Most reliable for a demo: scrcpy records the TV's picture and sound while the Mac's microphone records the room. I'd make a tvrec that runs both and merges them into one .mp4 with ffmpeg. Your voice comes from the Mac mic, and the answers come straight from the TV so they're clean.

The one thing in your voice that always ends up on the TV screen is the text: voice typing writes your question into the box.

Want me to install scrcpy and ffmpeg (brew install scrcpy ffmpeg), check that the answers are recorded, and make tvrec record both? You'd need to say a test question into the remote while I record.

You

yes, set it up

Agent

Ask Gemma is open on the TV. Next I start a scrcpy recording, have the TV ask "What is this?", and check whether the spoken answer was captured:

Sound was captured: silence for the first 9 s, then about 13 s of audio. I'll check a frame to confirm it's the answer being read:

Syntax checks out. Before handing it over, I'll test the merge step on the TV recording I already have, plus a stand-in mic track:

That caught a real bug: zsh reads $v[v] as an array lookup. I'm fixing it in tvrec and rerunning:

tvrec is set up. Open a new terminal (or run source ~/.zshrc) to pick it up.

tvrec            # TV picture + TV sound + your voice via the Mac mic → press Enter to stop
tvrec --no-mic   # TV picture and sound only

It saves ~/Desktop/Screenshots/tv-HHMMSS.mp4. While it records, the TV keeps playing sound normally, and there's no 3-minute limit any more.

What I tested:
- TV sound: recording a real question on the TV captured Gemma's spoken answer about the Shanghai photo, clearly, starting at 9 s.
- Merge step: I ran it on that recording plus a fake mic track and got a 21.7 s mp4 with the picture and sound lined up. This test caught a zsh quoting bug in the merge command, which I fixed.
- Not tested: a full run with your real voice through the Mac mic, since you're testing that yourself.

On your first run:
- Mic permission: macOS will ask whether your terminal app can use the microphone. Allow it, or your voice will be silent.
- Echo: the Mac mic also hears the TV's speakers, so the spoken answers may echo slightly over the clean copy. If that bothers you, I can turn the mic down while answers are playing, or switch to an external mic. Your mic list shows only the MacBook's built-in mic and a Zoom virtual device right now.

The new log alert is harmless. It's an Android warning from the video player as it shuts down: after a video was released, its stop() call reached a background thread that had already ended. The app didn't crash, and it isn't related to the Gemma changes. The video player code hasn't changed.

You

got this:

<pasted_content id="396a">
~ » tvrec me@mac
[2] 86157
[4] 86158
Recording… press Enter to stop.

[2] 86157 done scrcpy -s "$s" --no-window --no-control --audio-source=playback --audio-dup
[4] + 86158 exit 255 ffmpeg -nostdin -loglevel error -f avfoundation -i ":MacBook Pro Microphone"
Recording failed:
/opt/homebrew/Cellar/scrcpy/4.1_1/share/scrcpy/scrcpy-server: 1 file pushed, 0 skipped. 81.4 MB/s (733706 bytes in 0.009s)
[server] INFO: Device: [Sony] Sony BRAVIA 4K AE2 (Android 14)
scrcpy 4.1 https://github.com/Genymobile/scrcpy
INFO: ADB device found:
INFO: --> (tcpip) <tv-adb-serial> device BRAVIA_4K_AE2
</pasted_content id="396a">

Agent

Two separate failures here: the microphone recorder exited straight away (code 255), and scrcpy stopped before it ever began writing the video. I'll reproduce each on its own:

The mic part worked: your run saved about 18 s of mic audio. The problem is scrcpy. It connected to the TV and then never started recording. Reproducing that in an interactive shell like yours:

Testing the full path (TV plus mic) the same way your shell runs it:

That reproduces your failure. I suspect scrcpy only creates the file once the TV screen changes, and the screen was still both times. To check, I'll start a recording, wait, then press a key:

tvrec works now. Open a new terminal (or run source ~/.zshrc), then:

~ » tvrec
Press any button on the TV remote to start…
Recording… press Enter to stop.
Saving…
Saved ~/Desktop/Screenshots/tv-HHMMSS.mp4

Why your run failed: scrcpy doesn't start writing until something on the TV screen changes. Your screen was still, so it never started, and when you pressed Enter it quit with nothing saved. tvrec now waits for that first change, and pressing any button on the remote kicks it off. It also starts the mic at that moment, so your voice and the TV stay in sync. Separately, it now stops scrcpy with a signal the background process can't ignore; in my test it had ignored Ctrl-C and run on.

Echo: during tvrec the TV's speakers go quiet, so the Mac mic hears only you. Gemma's spoken answer still goes straight into the video, cleanly. You won't hear the answer in the room while recording, but it's in the file. tvrec --no-mic leaves the TV's speakers on.

Tested: I ran it the way your terminal does, with a simulated button press, and got a 7.6 s mp4 with the picture and the mixed audio lined up. I haven't been able to confirm by ear that the TV actually goes quiet, so check that on your first take.

I deleted my test recordings and the leftover temp files from your failed run. tv-131849.mp4 isn't one of mine as far as I can tell, so I left it. tv-130753.mp4 is my first silent test clip, so delete it if you don't want it.

You

there's a weird bug seems like, in the latest one there's a sound from the tv recorded even before the actual tv sound started, and it's overlapped with the actual tv sound

also no need to do fancy stuff for cancelling tv sound separwately, I have turned off tv sound while recording

Agent

I'll line up when the answer appears on screen with when the loud audio starts:

Before handing it back, I'll measure the sync. I'll press a key at a known moment and compare when the screen change shows up in the TV track with when the mic would have started:

My key press only moved the keyboard highlight, which is too small to measure. Trying again with Back, which closes the keyboard, a much bigger change on screen:

Fixed. Your voice and the TV now stay in sync in the recording.

The bug: scrcpy times its recording from the moment it connects to the TV, but tvrec only started the mic after the first change on screen. Your mic track started later yet was placed at 0:00 in the merge, so everything the mic heard came out several seconds early. In your recording, the mic also picked up the TV's speakers, which gave the early copy that overlapped the real, clean TV sound.

The fix: the mic now starts when scrcpy connects to the TV, so both tracks begin at the same moment.

  • Measured: I sent Back to the TV 4.16 s after the mic would have started, and the keyboard began closing at 4.14 s in the TV recording. That puts the two tracks within a fraction of a second of each other. This is a timing test; I didn't make a full recording with speech.
  • Simplified, as you asked: the TV's speakers stay on in every mode now. You mute the TV yourself to avoid echo.
  • "Press a button" step removed: recording starts right away. If you press Enter and nothing on the TV screen changed the whole time, it asks you to press a remote button first, because otherwise there'd be no video to save.

tv-132510.mp4 still has the offset audio, so you'll need to record that take again (open a new terminal or source ~/.zshrc first).

You

why do we need to understand a key press can't we directly start the recrding

Agent

You're right to question it. My explanation was probably wrong, and the step may not be needed at all. In the last test the TV sent video frames every 0.1 s from the very start, without any key press. So scrcpy was more likely waiting for audio: the capture mode I'm using (playback) sends nothing until some app plays a sound, and a key press makes a click sound. scrcpy has a second mode (output) that records everything the TV outputs, silence included, all the time. I'll test whether it starts writing straight away with no button press:

Neither mode writes anything to disk during the 6 s. But output still saved a complete 4.9 s recording when stopped, without any key press. Now the same check for the current playback mode:

Recording now starts right away, with no key press. But this clip came out 3.4 s long for a 6 s recording, so something is cutting the end off. Looking at the intermediate files:

No key press is needed. I was wrong about that: scrcpy was recording from the start the whole time. It keeps the data in memory and may only write the file when it stops, so the file looked empty while recording, and I misread that as "nothing is happening". A test with nothing happening on the TV still saved a full 5 s video with sound.

tvrec now just starts:

~ » tvrec
Recording… press Enter to stop.
Saving…
Saved ~/Desktop/Screenshots/tv-HHMMSS.mp4
  • Removed: the button-press check I'd added. A short check remains that only reports an error if scrcpy saved nothing.
  • Tested: with no key press. The TV and mic tracks both came out 3.6 s, so they start and stop together. I deleted the test clips.
  • Startup delay: there's still about 1–2 s before "Recording…" appears while scrcpy connects to the TV. Wait for that line before you start talking.
You

while recoridng somehow my mac seems to play the tv output sound, that seems like a bug csusing weird duplicates

Agent

That's scrcpy: --no-window hides the video window, but scrcpy still plays the TV's sound through the Mac's speakers by default, and the Mac mic was recording that. Adding --no-audio-playback:

I fixed it, and that was the cause of the duplicates. scrcpy's --no-window only hides the picture; it was still playing the TV's sound through the Mac's speakers, and the Mac mic recorded that as a second, slightly delayed copy. tvrec now runs scrcpy with --no-audio-playback, so the Mac stays silent and the TV's sound goes only into the file.

A test recording still saved the picture and TV sound fine. I didn't check by ear that the Mac stays quiet, since nothing was playing on the TV. Open a new terminal (or source ~/.zshrc) before your next take.

You

commit all the changes

Agent

I committed everything as 5176aa9 "Ask Gemma on a home server instead of Gemini in the cloud", on top of "Add Memories TV app". The working tree is clean.

The commit includes:
- Gemma instead of Gemini: Ask now uses Gemma through Ollama on your Mac. Settings has a server address and a model (default gemma4:e2b) in place of the Gemini API key.
- Speed: thinking is off and answers stream in as they're written, so a full answer takes about 2–3 s instead of 10–30 s.
- Plain HTTP allowed: Android otherwise blocks connections to Ollama on the home network.
- Ask panel: Down from the question field reaches the suggestions, and the hint is shorter.
- Top-bar buttons: they get the dark background so the labels read over bright photos.

tvrec isn't in the commit because it lives in your ~/.zshrc, outside the repo.

You
Agent

The repo is empty and private, so nothing gets overwritten. Pushing:

I pushed both commits to main on https://github.com/sbis04/memories_gemma. The repo was empty, so nothing was overwritten, and GitHub now shows 5176aa9 as the latest commit. I added the repo as the origin remote, and local main is set to track origin/main, so git push and git pull will work without extra arguments.

The repo is private. For the hackathon, judges will probably need to open it, so you'll likely want to make it public before you submit (GitHub → Settings → Change visibility).

You

this is the hackathon template can you fill it in a temporary .md file don't commit it

<pasted_content id="396a">
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

<!-- What does it do, and who is the friend or loved one you built it for? What problem does it solve for them? -->

Demo

<!-- Share a deployed link or a video demo. -->

Code

<!-- Show us the code! You can embed a GitHub repo directly into your post. -->

How I Built It

<!-- Which open-source AI did you use (open-weight models, agent harnesses, frameworks, local inference), and how is your project built around it? -->

Why Does Open Innovation Matter?

<!-- Why does open innovation matter for what you built? What did it make possible that a closed API wouldn't? -->

My Agent Session

<!-- Optional, but judges love it. Save your session with DevRelay and embed it with the agent_session tag (see the challenge page), or link to it. -->

Prize Categories

Best Use of Gemma
</pasted_content id="396a">

[Image #8]
[screenshot]

Agent

I'll draft it from what's actually in the app and what we measured. A few things only you know: who it's for, their reaction, and the demo video. Those get clearly marked TODO placeholders rather than invented details.

I filled in the template in SUBMISSION.md at the repo root. I also listed it in .git/info/exclude, a local-only ignore list, so git doesn't see it and it can't be committed by accident (git status is clean).

What's filled in, all from the code and the tests we ran:
- The app: what Memories does, and how Ask Gemma works (sending both the photo and the zoomed-in crop, the date/place/folder context, follow-up questions, streaming, reading aloud).
- The speed story: a table of the latency numbers we measured, the architecture diagram, and the three Mac setup commands.
- Why open matters: written around the actual switch from the Gemini API to local Gemma: privacy for family photos, no key or bill, nothing changing under you, and being able to tune it. It also says plainly that the trade-off is a computer at home having to be on.
- Code: the GitHub embed plus links to the two main files.

What only you can fill in (marked <!-- TODO --> in the file):
1. Who it's for, and the problem in your own words. This is the heart of "Build for a Friend", and writing quality counts most in judging, so it should be real.
2. The demo video. Upload a tvrec recording and embed it. I suggested a 60–90 s flow in the comment.
3. What they said when you handed it over. The challenge gives bonus points for this.
4. Your agent session embed, if you want to include it.

Before publishing: make the GitHub repo public, or the embed won't work for judges.

Want me to save this session to DevRelay so you have the {% agent_session %} embed ready?

Prize Categories

Best Use of Gemma

Top comments (0)