Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.
Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking to recommend a stack for this project, expecting to frown thinking that they still recommend tesseract. However, I found that they all recommend apple's vision framework. In fact the latest model to recommend Tesseract is gpt-4.1.
show comments
alt227
Slightly offtopic, but made me wonder.
Can you copyright things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?
EDIT: switched to the correct spelling of copyright.
show comments
lucideer
possibly off-topic, but for anyone interested in this on a more cross-platform / holistic basis, Immich does this
(& by "this" I mean an approximate AI search for photos & videos - I can't account for the "every frame", nor for the comparative search quality)
show comments
measure2xcut1x
How well do you think this would work on stock photography on m1 mac with 32GB ram? For example I'd like to be able to search a folder of ~2k photos for houses with palm trees. Or find photos of kitchens, or find photos of desert southwest landscapes.
show comments
yt1998
Why chose CLIP to do this. Have you tried small VLMs like Qwen-VL? I believe those models have video encoders can better perform at this scenario.
show comments
hn3ufz62f7
Having built something similar with CLIP on an M1, frame sampling rate is the whole ballgame. One frame a second on 12k videos is days, keyframes only got me to an overnight run.
show comments
stephenitis
I like the entire premise, the one thing stopping me from trying this is not knowing the time scales that I will need to set my computer aside for the processing of large folders of video frames, or my photos library's videos, some 12,000 videos
qlasisi15
This would really help in video editing.
Thanks!
show comments
hemedanmert
This is so cool! it could be really beneficial for editors
isnt it expensive tho?
show comments
cpursley
Why is this a JS bloatware instead of native or Rust which is easier than ever now with LLM coding tools.
show comments
qprofyeh
This is cool. Any way to search for People / faces / pets? Like on iOS?
aavisangle
How do you do that? What's the architecture? Can you guide on that?
freecodeio
would be lovely if picture embeddings were attached to the file by the camera but one can only dream of such futures
show comments
Jeeetendra
jumping straight to the right moment in a video is the useful bit here. how often does the balanced sampling miss something that only appears for a second or two?
doubleorseven
how long does it takes for a FHD 90 minutes asset?
mannanj
its a cool project, but I dont want to consume someones ai slop to discern whats true. if the human wrote the page in their own words, I would have considered using it.
otherwise, I can just make my own with my own ai. why consume someones slop when I can eat my own.
Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.
Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking to recommend a stack for this project, expecting to frown thinking that they still recommend tesseract. However, I found that they all recommend apple's vision framework. In fact the latest model to recommend Tesseract is gpt-4.1.
Slightly offtopic, but made me wonder.
Can you copyright things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?
EDIT: switched to the correct spelling of copyright.
possibly off-topic, but for anyone interested in this on a more cross-platform / holistic basis, Immich does this
(& by "this" I mean an approximate AI search for photos & videos - I can't account for the "every frame", nor for the comparative search quality)
How well do you think this would work on stock photography on m1 mac with 32GB ram? For example I'd like to be able to search a folder of ~2k photos for houses with palm trees. Or find photos of kitchens, or find photos of desert southwest landscapes.
Why chose CLIP to do this. Have you tried small VLMs like Qwen-VL? I believe those models have video encoders can better perform at this scenario.
Having built something similar with CLIP on an M1, frame sampling rate is the whole ballgame. One frame a second on 12k videos is days, keyframes only got me to an overnight run.
I like the entire premise, the one thing stopping me from trying this is not knowing the time scales that I will need to set my computer aside for the processing of large folders of video frames, or my photos library's videos, some 12,000 videos
This would really help in video editing. Thanks!
This is so cool! it could be really beneficial for editors
isnt it expensive tho?
Why is this a JS bloatware instead of native or Rust which is easier than ever now with LLM coding tools.
This is cool. Any way to search for People / faces / pets? Like on iOS?
How do you do that? What's the architecture? Can you guide on that?
would be lovely if picture embeddings were attached to the file by the camera but one can only dream of such futures
jumping straight to the right moment in a video is the useful bit here. how often does the balanced sampling miss something that only appears for a second or two?
how long does it takes for a FHD 90 minutes asset?
its a cool project, but I dont want to consume someones ai slop to discern whats true. if the human wrote the page in their own words, I would have considered using it.
otherwise, I can just make my own with my own ai. why consume someones slop when I can eat my own.