I work on OpenSubs, a free, open source (AGPL-3.0) subtitle tool that runs entirely in the browser tab.

You drop in a video file, and Whisper transcribes it on your own machine, using transformers.js with WebGPU where available and WebAssembly otherwise. The model (40–250 MB) downloads once and is cached. There is no upload endpoint in the product, so the video has nowhere to go.

After that you can:

  • fix lines by typing over them (click a timestamp to jump to that moment)
  • translate into 20 languages with Chrome's built-in on-device translator
  • pick one of 12 caption styles, including word-by-word highlighting
  • export SRT / VTT / ASS, or burn the subtitles into an MP4 (libass compiled to WebAssembly, encoded with WebCodecs)

A few things I learned building it:

  • Whisper hallucinates on silence and music ("Thanks for watching!", or the Japanese equivalent). A Silero VAD pass runs before Whisper, and a cleanup step drops the known stock phrases.
  • Singing doesn't count as speech for the VAD, so a music video gets a "no speech found" warning. You can still force it.

Honest limits: it only takes video files, not audio-only files. Cue timings can't be edited yet. Builds are release candidates. Everything that runs locally is free with no account; the only paid part is optional cloud translation on our backend (US$5 for 1000 credits), and you can bring your own Claude / OpenAI / DeepL key instead.

Site: https://opensubs.app/ Code: https://github.com/open-subs/opensubs

Feedback welcome, especially on languages where the transcription goes wrong.

I just made an open-shoe! This is the first FOSS shoe ever. Everything is free but if you want both shoes, you'll need a subscription. OpenShoe is right foot only. Any foot, any size, durable and incredibly versatile. Left shoe only available in sizes M5, M7, M11.5, and M16.

Most videos I watch are in English, Spanish or German, I don't need an translation for it, less one by AI.

Nice, another blindly generated app by Claude Code. 37 commits and already announcing it tells it all.

yep, I would rather generate one with Claude to my liking and make it more efficient (target the CLI for example)

Which is absolutely legit. I only criticize the people releasing a “production ready” product, which is nothing but a PoC and then also claim that it is AGPL-3, while abandoning it after 2 months. This hurts the open source community imho.

while abandoning it after 2 months. This hurts the open source community imho.

I dont expect them to maintain it for eternity. The problem is people expect too much from the authors as if it's the author's duty to fix bugs and add features. And tbh, this attitude hurts open source community more than anything. Ffs, the code is open source, do what you like. The author is not responsible for everything. Be grateful that the code is open-source.

I am not talking about the many projects, where you have a single developer just enjoying making software, while providing it to the open source community. I have zero expectations to those projects.

I am talking about people promoting their open source products. Making a fancy website. Act professionally. All this behavior causes people to have expectations. So why invest time into promoting it in the first place? There is literally no need to do this, expect for the owner to attract more and more people.

Don’t promise a sky scraper if you can only maintain a bungalow.

This hurts the open source community imho.

Nah. At worst the open source community is indifferent. If someone can pick up where this left off then it benefits the community.

All "generated code" hurts open source, actually.

Badly written code hurts open source, but humans are just as guilty of that as machines.

How? Genuinely curious.

Whisper hallucinates on silence and music (“Thanks for watching!”, or the Japanese equivalent).

I've had Whisper ask me to donate to anime fansubbing groups. I wonder what it was trained on?

Sometimes if I hit the voice typing on my phone, which I'm pretty sure is whisper, and then don't say anything, it just comes out with nonsense.

Just did this and also got...

Thanks for watching!

It sounds like a pretty cool project, thanks for sharing! As a browser-based project designed to run offline, you might want to consider shipping it as an electron app.

tauri is better and more efficient for this.

Why not just use the whisper plugin for Bazarr?

i get ya, but also a lot of people don't have a dedicated self hosted setup. recently my hdd died and all my films were there, so now i'm downloading and deleting, which just feels like a drag to start bazarr for one film. this is handy for some people.

What’s the max file size?

I’ve been looking for the fan edit of marvels Infinity saga.

It’s about 50 hours long divided into many many files ranging from 3-10 gb each. Would it be able to handle this? Or what’s the best way to split them up

I'm a lazy reader so don't know if this is implemented but it would be cool if it could be done live with the current audio out of the system. I know this was done for streamers etc using OBS/localvocal but it's a pain to get working in my experience - having to fiddle around with OBS filter settings.

Open source, but made with a copyright infringing LLM.

oh no, how awful of them to release the code back into the wild for free with a copyleft license. truly an awful thing to do.

i love how anti ai people have become the wardens of copyright somehow.

Statement of fact. Claude was engineered with copyright material.

whats wrong is not copyright theft. whats wrong is labour that was stolen.

Yes and no. As an author of open source projects, they were made available for people to fork, improve and conrribute under permissive licenses. Then come LLMs that uses code like mine to generate code or slop. My choice about license was made before LLMs. Now they use derivatives of my work without attribution as required by the license.

Even if I change the license, it won't be respected, and nothing else will change. To me, that's copyright infringement on me, ignoring the hours I put in.

On a separate note, LLMs only produce more of the same, ignoring the environmental and economic impacts. Often, more of the same is not what we need. It leads to the same mistakes being made. I've never used LLM to code, never will.

as a pirate, i shit on copyright, culture should be free. the issue with llm is not that they used copyrighted shit, if the llm would be open source and not owned by a corpo, i'd be contributing to knowledge. here, you're just contributing to some gringo westoid billionaire's pockets. everything is a remix, i dont have an issue with that. so again, the issue is labour theft in marxists terms.

Whisper is not an llm, just because it is trained on big datasets does not make it an LLM. Copyright infringement still happened during the training.

no, but claude is, and it's the main contributor to the repo.

Which software are we talking about? Unfortunately no one links to it, so I will guess. Whisper: website | source is an Ai model, specifically powered by OpenAi.

Whisper is not an llm, just because it is trained on big datasets does not make it an LLM.

The training dataset consists of 680,000 hours of labeled audio-transcript pairs sourced from the internet using semi-supervised learning.

Brother, I don't think you know what a large language model is if you think it being trained on a human lifespan's worth of audio transcripts isn't enough.

Its job is NLP. It is an LLM, if not as general-use as something like ChatGPT.

LLM is a subset of NLP. All LLMs are NLP but not all NLP are LLMs. To be more precise, LLM is a method of achieving NLP while NLP is the goal. Whisper is just a specialised speech-to-text engine. It doesn't have the text reasoning engine that is the core of LLMs. Yes, it has a dataset to inform it's encoder and decoder, but that doesn't make it an LLM. If you strip out the encoder, then you're getting closer.

If you strip out the encoder, then you're getting closer.

That's flat-out not true. BERTs, for example, are LLMs and are encoder-only, just like GPTs are decoder-only. Whisper is trained on a large corpus of text and does NLP tasks. That's all it is to be an LLM.

A transformer model being encoder-only, encoder-decoder, or decoder-only has fuck-all to do with whether it's an LLM. In fact, it being a transformer model at all isn't required for it to be an LLM.

BERT is a special case. The academic consensus is rightly considering it an LLM since it operates on text but the industry doesn't consider it an LLM since LLM in its colloquial meaning has drifted to mean a model generating text (which BERT doesn't). Whisper is not an LLM since it's working on audio, so it is an Automatic Speech Recognition model (ASR).

Text in = LLM in academic circles (BERT and GPT)

Text in, text out = LLM in all circles (GPT only)

Audio in, text out ≠ LLM (Whisper)

Now, we can argue whether we want to accept the colloquial meaning of LLM, but the fact is that Whisper is not an LLM. And neither ASR nor BERT causes even a fraction of the damage GPT is doing, but that's another discussion.

@opensubs

In the browser

Never gets uploaded

What are you talking about?

The browser runs a lot of stuff that doesn't get passed on to the servers you connect to. In this case it can run a local whisper model locally and pass your video stream through it, rather than pass it through some server online.

@virku Then what's the point? You can just download the model and run it locally without using this service...

Ease of use..

It's not a "service" but a program that runs locally.

Downloading the model and running it locally is essentially what it does. If you do it by yourself then you'd have to write your own scripts. At which point, you might as well share those scripts with the community putting them in Github, and maybe including them in a web extension to integrate with the browser and/or in a static website that locally runs it via WASM... precisely what this project is doing...

This is how loads of things run these days. Browser handles the UI because it's way easier to get working on all kinds of different user machines. I don't like it personally but that's why.

It's how pretty much any server dashboard runs.

If i understood this right...

Instead of uploading a file you just download the model. The sound then gets transcribed as usual without the service accessing the file.

@beyoublahaj So audio gets uploaded?

oh nah I reread the post and edited my comment. I misunderstood too ^^`

@beyoublahaj So what's the point of this service if you can download the model yourself and use it without this website?..

As people have tried to explain - it's not a website. The UI for the software simply runs via the browser.

You don't have to use the website.. the project is not just that website, it's also packaged as a browser extension, a desktop app and a CLI program.

That website is just a convenient way to see it working (while still processing it locally) without installing anything.

If you download the raw model and write up scripts to create subtitles out of video by yourself then you are essentially re-building the same app.. the result would be similar, just duplicating effort.

I assume you can use it on any platform that runs a browser, without installing software on the system. I use (other type of) software that runs in the browser, but locally. This type of browser software is handy.

No idea, I suppose that it's more convenient and easier.

midwest.social

Rules

  1. No porn.
  2. No bigotry, hate speech.
  3. No ads / spamming.
  4. No conspiracies / QAnon / antivaxx sentiment
  5. No zionists
  6. No fascists

Chat Room

Matrix chat room: https://matrix.to/#/#midwestsociallemmy:matrix.org

Communities

Communities from our friends:

Donations

LiberaPay link: https://liberapay.com/seahorse