OCR Subtitle Ultra is a practical utility designed to extract hard-coded subtitles from videos using Optical Character Recognition (OCR). It automatically converts on-screen text into standard SRT files, saving hours of manual transcription for editors, educators, and content creators.
Driven by an optimized layout-parsing architecture and a high-performance multimedia backend, version 3.0 marks a massive leap forward—introducing intelligent auto-detection workflows while retaining full manual precision controls for professional creators.
(Note: Recognition accuracy is subject to video resolution, font styles, and background contrast.)
【What is New in v3.0: Automation & AI Workflows】
Intelligent Auto-Detection with Manual Controls
Version 3.0 features a smarter layout engine that automatically detects the subtitle region (ROI), locks text bounding boxes, and auto-identifies the video's spoken language upon import for a zero-configuration experience. Professional control is never compromised: users can seamlessly switch to manual mode to custom-adjust the scanning area or specify languages for highly complex, non-standard video layouts.
Seamless AI Subtitle Translation Workflow
Unlock global localization instantly. The app intelligently packages your extracted SRT text with tailored translation prompts. On iOS, leverage the native Share Sheet to route text directly to your preferred AI assistants (such as ChatGPT or Claude). On macOS, the app neatly formats text into your system clipboard for a frictionless paste-and-go experience into any desktop AI or web assistant.
AI Video Copilot (Instant Video Q&A)
Productivity shouldn't end with an SRT file. With our built-in high-frequency AI prompt templates, you can instantly bundle your transcript and send it to an AI to perform advanced tasks:
Fact-Checking: Verify claims or detect clickbait within the video content.
Instant Summary: Condense long-duration footage into structured bullet points.
Copywriting & Titles: Generate high-CTR titles, video descriptions, and marketing hooks.
Thumbnail Inspiration: Brainstorm detailed visual prompts to design compelling video cover art.
Integrated Standard Format Video Previewer
Never worry about how to play or mount your newly generated SRT files. Version 3.0 features an integrated video preview player. It seamlessly supports system-compatible standard formats (such as MP4, MOV, M4V), allowing you to play your media with the newly created subtitle track overlaid right inside the app for an instant preview experience.
【Core Performance & Privacy Security】
Strict "Zero-Storage, Zero-Copy" Architecture: Your data security is our highest priority. Operating strictly within the Apple Sandbox framework, the application processes video files entirely on-the-fly. It never duplicates, caches, or permanently stores your video files in local containers or remote servers.
Temporary Access Lifecycle Management: The application securely requests transient, read-only media access via system-secured tokens (Security-Scoped URLs) in memory. Once the security token expires, the resource is automatically revoked. Your media remains entirely under your own control.
Robust Codec Support: Built on a rigorously compliant multimedia backend, the app delivers ultra-stable decoding for modern codecs including H.265 (HEVC), AV1, VP9, and more. Advanced memory management prevents crashes, ensuring flawless processing for 4K Ultra-HD footage and long films.
Universal Workflow Integration: Exported standard SRT files are fully editable and 100% compatible with YouTube, VLC, Final Cut Pro, Adobe Premiere, DaVinci Resolve, and major editing platforms.
【Support & Feedback】
I am deeply committed to continuously refining OCR Subtitle Ultra based on real-world production environments and user feedback. If you encounter specific video encoding edge cases or have feature suggestions, please contact me directly:
Email: hanmingjie@gmail.com
Web: www.hanmingjie.com
I'm very satisfied with both the OCR accuracy and processing speed. I've tried several OCR applications, and this one stands out for its excellent recognition accuracy and fast performance, which has significantly improved my workflow.One feature I would love to see is a job queue. Currently, the app can process only one video at a time. It would be much more convenient if multiple videos could be added to a job queue and processed automatically one after another. It would also be great if each queued job could use its own template, allowing videos that require different templates to be queued together and processed automatically in sequence.With this feature, I believe the app would be close to perfect. I look forward to seeing this app continue to improve with future updates. Thank you for creating such a great application!
“性能不错,但双行字幕存在致命问题”
"Good performance, but critical issue with two-line subtitles"
Jungwe2
性能不错,但在双行字幕中无法指定提取哪一行。它只能识别其中一行,这是一个致命的问题。请允许用户手动指定要提取的字幕范围,而不是仅依赖自动检测。The performance is good, but it doesn't allow you to specify which line to extract in two-line subtitles. It only recognizes one of the two lines, which is a critical flaw. Please allow users to manually specify the subtitle range instead of relying solely on automatic detection.
개발자 답변
非常感谢您的深度反馈!您提到的手动指定字幕范围功能确实非常关键。正如您所期待的,我们在 1.9 版本中升级了多行字幕的处理逻辑,现在您可以更加精准地控制提取范围,不再受限于自动检测。诚邀您观看我们的演示视频了解新功能的操作方式:https://www.youtube.com/watch?v=l9z-OVaNOzs。希望这次更新能显著提升您的效率,期待您的再次评价!Great news! Based on your suggestion, we’ve upgraded our multi-line subtitle scanning in v1.9. You can now manually choose the target line for extraction.We truly appreciate your professional feedback, which has helped us continue to improve. In version 1.9, we have introduced a more flexible recognition mechanism that solves the issue of being unable to specify a specific line.Check out the demo here: https://www.youtube.com/watch?v=l9z-OVaNOzsThank you for using our app on Mac Catalyst! If this update resolves your concern, we would be grateful if you could consider updating your rating.
1. New Whisper Voice Recognition Mode
Integrated the Whisper speech-to-text framework. When videos lack on-screen text or have highly cluttered backgrounds, users can select this mode to directly transcribe audio tracks into time-coded standard SRT files. Supports automatic audio language detection.
2. Refactored Functional Modes (Four Independent Options)
To better accommodate varied post-production workflows, core features are now structured into four distinct, mutually exclusive modes:
Manual ROI Mode: Users manually define the Region of Interest (ROI) for targeted visual OCR scanning.
Auto ROI Mode: The system automatically detects and locks the subtitle area for visual OCR scanning.
Voice Recognition Mode: A new mode that bypasses visual scanning entirely, performing voice-to-text transcription directly on the audio stream.
Playback Only Mode: A new mode that invokes neither OCR nor voice recognition engines. It functions strictly as a standard video player, allowing users to preview media or overlay existing subtitle tracks.
3. Core Workflow Integration
Subtitles generated via either OCR scanning or the new voice recognition mode integrate seamlessly with the pre-existing Seamless AI Translation Workflow and AI Video Copilot prompt templates.
4. Sandbox Privacy & Performance Optimization
Strictly operates within the Apple Sandbox framework. Both audio and video streams are processed on-the-fly inside active memory with zero local copying, caching, or remote server exposure. Optimized dynamic memory release logic to improve stability when handling long-duration footage.
버전 3.1
韩 明洁 개발자가 아래 설명된 데이터 처리 방식이 앱의 개인정보 처리방침에 포함되어 있을 수 있다고 표시했습니다. 자세한 내용은 개발자의 개인정보 처리방침 을 참조하십시오.
데이터가 수집되지 않음
개발자가 이 앱에서 데이터를 수집하지 않습니다.
개인정보 처리방침은 사용하는 기능이나 사용자의 나이 등에 따라 달라질 수 있습니다. 더 알아보기