I experimented with this exact same approach earlier this year.
It's barely sufficient, because, bluntly, most apps just aren't wired up right.
So you end up having to hand code a lot of specific profiles for specific apps to make this work well, and even then, you don't quite get the right level of detail to make it work out.
Will try this app, to see if it improved on my own approach, but man, the hope levels are low.
One app that's using this technique (not exactly sure if it's the same) is Littlebird: https://littlebird.ai/
I also saw that HeyClicky started doing something similar but end up removing from the product.
> reads the text of your focused window every few seconds through the Accessibility API
> It writes plain markdown
Where are the formatting decisions coming from?
neat. why not screenshot and tesseract (videos/images/viewport/etc)
Because you then have the macOS orange screen sharing warning/icon. I don't really want to record my screen, just the text is enough.
Agree
I experimented with this exact same approach earlier this year.
It's barely sufficient, because, bluntly, most apps just aren't wired up right.
So you end up having to hand code a lot of specific profiles for specific apps to make this work well, and even then, you don't quite get the right level of detail to make it work out.
Will try this app, to see if it improved on my own approach, but man, the hope levels are low.
I haven't read about how the Codex Appshots work yet, but this can be used to extract text properly. I guess. How does this idea look to you?
One of the reasons people like TUIs is because the text is always just right there.
Not if it's rendered on GPU, I guess?
Yeah don't get your hope up too much. I need to push through more cleaning up of what's captured. Let me know how it goes, keen for some feedback