Syncing a YouTube Caption Bar with Video Time and Binary Search
Making captions larger sounds like a CSS task. The difficult part begins after the bar exists: it must follow playback, react immediately after a seek, survive YouTube's single-page navigation, and avoid rewriting the DOM hundreds of times for the same sentence.
This article describes the synchronization core of YouTube Big Caption Bar, a Manifest V3 Chrome extension that places a large reading bar above a YouTube video. The design uses the media element as its clock, a sorted transcript as its timeline, binary search as its matcher, and a small amount of state to prevent unnecessary rendering.
The short version
Read time from the video, normalize cues into one sorted array, locate the active cue with binary search, and touch the DOM only when that cue changes. Reset the whole loop when YouTube navigates to another video.
Treat the video element as the source of truth
Do not build a separate stopwatch. Playback can pause, buffer, change speed, or jump after the user seeks. A timer that merely adds elapsed wall-clock time will drift from the actual media.
HTMLMediaElement.currentTime reports the current playback position in seconds. Assigning a value to it also seeks the media, so the same clock can later support click-to-seek interactions. See the MDN currentTime reference for the platform behavior.
function findVideo() {
return document.querySelector("video.html5-main-video") ??
document.querySelector("video");
}
function readPlaybackTime(video) {
return Number.isFinite(video?.currentTime) ? video.currentTime : 0;
}
The YouTube-specific selector is useful, but the generic fallback makes the code less brittle. The extension should still handle a missing element because YouTube may be between route states when the content script runs.
Normalize the transcript into one contract
Synchronization becomes simple when every transcript provider returns the same sorted shape.
[
{ start: 0.42, text: "Welcome back", translation: "" },
{ start: 2.90, text: "Today we will build...", translation: "" }
]
The UI should not know whether entries came from test data, a supported caption source, or another provider. It only needs ascending start values and safe strings.
function normalizeTranscript(items) {
return items
.map((item) => ({
start: Number(item.start),
text: String(item.text ?? "").trim(),
translation: String(item.translation ?? "").trim()
}))
.filter((item) => Number.isFinite(item.start) && item.text)
.sort((a, b) => a.start - b.start);
}
This boundary is also a good place to remove empty cues and reject malformed timestamps. Without normalization, a single NaN or out-of-order entry can make a correct search algorithm appear unreliable.
Match the active cue with binary search
For a cue at index i, it is active when playback has reached its start and has not yet reached the next cue's start.
cue[i].start <= currentTime < cue[i + 1].start
The last cue remains active after its start unless the transcript includes an explicit end time. A binary search finds that interval in logarithmic time.
function getCurrentCaption(transcript, currentTime) {
let low = 0;
let high = transcript.length - 1;
while (low <= high) {
const mid = Math.floor((low + high) / 2);
const current = transcript[mid];
const next = transcript[mid + 1];
if (currentTime < current.start) {
high = mid - 1;
} else if (next && currentTime >= next.start) {
low = mid + 1;
} else {
return current;
}
}
return null;
}
A linear scan can be acceptable for a short clip, but it repeats work from the beginning on every tick. Binary search scales cleanly for long lectures and transcripts with thousands of cues.
There is one product decision hidden here: gaps. If cues have only start times, the previous sentence stays visible until the next one begins. If the source provides explicit durations, add an end field and return null when currentTime >= current.end.
Choose a synchronization trigger
The browser fires a timeupdate event when the media time changes, but its frequency varies with system load. MDN notes that user agents may deliver it roughly between 4 and 66 times per second, so code should not assume a fixed cadence. See the MDN timeupdate reference.
There are three practical strategies:
| Strategy | Strength | Tradeoff |
|---|---|---|
timeupdate |
Event-driven and simple | Frequency is browser-controlled |
| Fixed interval | Predictable upper bound on work | Wakes while paused unless managed |
requestAnimationFrame |
Smooth visual updates | Usually more work than sentence cues need |
For sentence-level captions, a 250–300 ms interval is usually responsive enough. Add immediate synchronization on seeked and play so a jump does not wait for the next tick.
let intervalId = null;
function startSync(video, sync) {
stopSync();
intervalId = setInterval(sync, 300);
video.addEventListener("seeked", sync);
video.addEventListener("play", sync);
sync();
}
function stopSync() {
if (intervalId !== null) {
clearInterval(intervalId);
intervalId = null;
}
}
Production code should retain the listener function and remove it when switching videos. Otherwise repeated route changes can accumulate listeners even though the interval itself is cleared.
Render only when the cue changes
The matcher may run several times while the same caption is active. Updating textContent on every tick is unnecessary. Keep the active cue or index in state and render only on a transition.
let currentCaption = null;
function syncCaption(video, transcript, elements) {
const caption = getCurrentCaption(transcript, video.currentTime);
if (caption === currentCaption) return;
currentCaption = caption;
elements.text.textContent = caption?.text ?? "";
elements.translation.textContent = caption?.translation ?? "";
}
Using textContent rather than innerHTML is important because caption text is external data. It prevents subtitle markup from becoming executable page content.
Object identity works if the normalized transcript array stays unchanged. If entries are recreated during updates, store the active index or a stable cue ID instead.
Survive YouTube's single-page navigation
YouTube changes videos without performing a traditional full-page load. A content script registered for https://www.youtube.com/* can therefore remain alive while the URL, video element, and transcript all change.
Chrome content scripts run in an isolated world by default, which keeps extension JavaScript separate from the page's own variables while still allowing DOM access. The registration model is documented in Chrome's content script reference.
On navigation, the extension should:
- Read the new
vquery parameter. - Ignore the event if the video ID did not change.
- Stop the old synchronization loop.
- Clear the active cue and transcript.
- Find the new
<video>element. - Load and normalize the new transcript.
- Start one synchronization loop.
let currentVideoId = null;
document.addEventListener("yt-navigate-finish", () => {
const videoId = new URL(location.href).searchParams.get("v");
if (videoId === currentVideoId) return;
currentVideoId = videoId;
stopSync();
currentCaption = null;
// Find the new video and load its transcript here.
});
YouTube-specific DOM and custom events are implementation details, not a stable public API. Keep them behind small adapter functions and provide timeouts or fallback observers so a markup change does not break the entire extension.
Separate page work from network work
The content script is the right place to find the video, read currentTime, and render the bar. Network work and extension-level coordination belong in a separate component such as the Manifest V3 service worker.
// content.js
const response = await chrome.runtime.sendMessage({
type: "GET_CAPTIONS",
videoId
});
Chrome supports one-time messages between content scripts, extension pages, and service workers through runtime.sendMessage(). The official messaging guide describes the available patterns.
Do not make the synchronization engine depend on one undocumented extraction technique. YouTube page internals and caption endpoints can change. A provider interface lets the UI remain stable while caption acquisition is replaced, repaired, or disabled.
Test the timeline boundaries
The matcher is small enough to test without a browser.
const cues = [
{ start: 1, text: "A" },
{ start: 3, text: "B" },
{ start: 8, text: "C" }
];
console.assert(getCurrentCaption(cues, 0.99) === null);
console.assert(getCurrentCaption(cues, 1).text === "A");
console.assert(getCurrentCaption(cues, 2.999).text === "A");
console.assert(getCurrentCaption(cues, 3).text === "B");
console.assert(getCurrentCaption(cues, 20).text === "C");
Browser-level tests should add seeking backward, changing playback speed, pausing for several minutes, navigating to another video, opening a non-video route, and toggling the bar off and on. Also verify that only one timer and one set of event listeners remain after repeated navigation.
Conclusion
Reliable caption synchronization comes from choosing the right boundaries. The video element owns time. Providers own transcript acquisition. A normalized sorted array owns the timeline. Binary search selects the active cue, and the renderer changes the DOM only when that cue changes.
This structure is modest, but it handles the cases that make a caption overlay feel dependable: pause, seek, long transcripts, route changes, and source failures. Once the clock and matching layer are stable, font controls, translation display, bookmarks, and click-to-seek can be added without rewriting the core.
Comments
Post a Comment