DEV Community

liveavabot
liveavabot

Posted on

Why iPhone Videos Silently Fail as Telegram Avatars

The problem nobody talks about

You record a short clip on your iPhone, open Telegram, tap "set video avatar", upload, and nothing happens. No error, no toast, just the old photo sitting there. I hit this once before I got curious.

The culprit is HEVC. Since iOS 11 the default camera codec is H.265 (HEVC), and Telegram's video avatar pipeline rejects HEVC without telling you. The upload "succeeds" because the file reaches the server, but the server side validator throws it out.

What Telegram actually wants

The video avatar spec is strict and not well documented. After some trial and error on what gets accepted, here is the shortlist:

  • Codec: H.264 (AVC), profile baseline or main
  • Pixel format: yuv420p, not yuv420p10le
  • Container: MP4 with faststart
  • Resolution: 800x800, square
  • Duration: 1 to 10 seconds
  • File size: under 2 MB
  • Audio: must be absent, not just muted

Miss any of these and the avatar silently reverts. The "under 2 MB" one is brutal: a 9 second 800x800 clip at reasonable bitrate often lands at 3 MB and gets dropped.

FFmpeg does the heavy lifting

The pipeline I settled on has two passes. First, cropdetect to find the content bounding box on vertical phone video so I can square crop without black bars. Second, scale and encode.

# Pass 1: detect crop box over 2 seconds
ffmpeg -hide_banner -ss 0 -t 2 -i input.mov \
  -vf "cropdetect=24:16:0" -f null - 2>&1 \
  | grep -o "crop=[0-9:]*" | tail -1
# -> crop=1080:1080:0:420

# Pass 2: encode to Telegram safe MP4
ffmpeg -y -i input.mov -t 9.5 \
  -vf "crop=1080:1080:0:420,scale=800:800:flags=lanczos,format=yuv420p" \
  -c:v libx264 -profile:v main -preset slow -crf 28 \
  -movflags +faststart -an output.mp4
Enter fullscreen mode Exit fullscreen mode

A few things worth calling out:

  • -an strips audio entirely. Mute is not enough, the stream has to be gone.
  • -movflags +faststart moves the moov atom to the front so the server can validate without downloading the whole file.
  • -crf 28 sits in a sweet spot for 800x800 clips: visibly fine, usually under 2 MB for 8 to 10 second content. For longer clips drop to 30 or trim duration.
  • lanczos for scaling because bilinear on faces looks mushy.

If the output still exceeds 2 MB after this, I either shorten the clip or bump CRF to 30 or 32 and re-encode. That retry loop runs once in maybe 15 percent of uploads.

Aiogram 3 handler

Here is a minimal handler that catches an incoming video or document, pipes it through the ffmpeg converter, and sends the result back. Error handling trimmed for brevity.

from aiogram import Router, F
from aiogram.types import Message, FSInputFile
from pathlib import Path
import asyncio
import tempfile

router = Router()

@router.message(F.video | F.document | F.animation)
async def handle_video(message: Message) -> None:
    file = message.video or message.document or message.animation
    if not file:
        return

    with tempfile.TemporaryDirectory() as tmp:
        src = Path(tmp) / "input.mov"
        dst = Path(tmp) / "output.mp4"

        tg_file = await message.bot.get_file(file.file_id)
        await message.bot.download_file(tg_file.file_path, src)

        ok = await convert_to_avatar(src, dst)
        if not ok:
            await message.answer("Could not convert this clip. Try something shorter.")
            return

        await message.answer_video(
            FSInputFile(dst),
            caption="Set this as your video avatar in Settings, Edit Profile.",
        )

async def convert_to_avatar(src: Path, dst: Path) -> bool:
    crop = await detect_crop(src)
    cmd = [
        "ffmpeg", "-y", "-i", str(src), "-t", "9.5",
        "-vf", f"{crop},scale=800:800:flags=lanczos,format=yuv420p",
        "-c:v", "libx264", "-profile:v", "main", "-preset", "slow",
        "-crf", "28", "-movflags", "+faststart", "-an", str(dst),
    ]
    proc = await asyncio.create_subprocess_exec(
        *cmd,
        stdout=asyncio.subprocess.DEVNULL,
        stderr=asyncio.subprocess.DEVNULL,
    )
    await proc.wait()
    return (
        proc.returncode == 0
        and dst.exists()
        and dst.stat().st_size < 2 * 1024 * 1024
    )
Enter fullscreen mode Exit fullscreen mode

The real bot does a bit more: size check retry loop, cropdetect fallback for very dark clips, user facing progress on long files. The skeleton above is what makes the thing work end to end.

Packaging it as a bot

I wrapped this into a Telegram bot and let it sit on a small VPS. Send any clip, get back a video avatar ready to upload. It handles HEVC from iPhones, portrait from Android, square clips, GIFs, even old DV footage from a mini-DV tape I digitized recently (that was fun).

Try it: @LiveAvaBot

If you just want the ffmpeg one liner for your own scripts, the two commands above are the whole product. The bot wrapper is convenience.

Lessons and edges

A few things I did not expect going in:

  • iPhone portrait video has a rotation flag, not a rotated raster. ffmpeg honors the flag by default on recent builds, older builds need -vf transpose=1 or the output is sideways. Pin a known ffmpeg version in your Docker image.
  • GIFs loop by default inside Telegram, so I keep looping behavior intact in the MP4 by not setting an audio track and letting the player decide.
  • Users sometimes send 4K ProRes from pro cameras. Those files are 400 MB. I reject anything over 50 MB before download instead of waiting through a slow upload to the bot.
  • Animated WebP from stickers is a trap. Treat it as an animation but convert the frame rate explicitly, otherwise you get a static frame.

Next thing I want to add is a face centered square crop for portrait clips where the subject sits in the bottom third. Current heuristic uses cropdetect only, which centers on content density, not faces. Probably mediapipe or a tiny ONNX model, still deciding.

Built by me, @liveavabot.

Top comments (0)