Ironically, the pill counting example selected to showcase "the best vision model" can be easily solved with OpenCV template matching, a technology created 25 years ago.
Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.
My exposure to Claude-produced UIs is limited, but I have started to notice certain design trends they tend to have in-common, which might be becoming hallmarks of AI-produced UIs - the same way we've started noticing the clichés of low-effort LLM-generated text.
FWIW, the summary-description[1] of "frontend-design"[2] gives me a few things to pick at:
> create polished code
Methinks only if you're using it with a very popular framework like React. What happens if you ask Claude to make the UI in WinForms or MFC?
> high-impact animations
That's bad UX 101 right there: animations in a UI exist as an affordance to the user, and never for its own sake (e.g. macOS's "genie" animation when you minimize a window to the dock exists so the user knows where they can restore the window from). The only people who actually want "high impact animations" in software are salespeople who want something for demo purposes.
> generic system fonts, predictable purple gradients, and cookie-cutter components.
This screams wanting to be different for the sake of standing-out, not because it results in a better software product; users benefit when their software fits-in with platform conventions: if you refuse to use a stock checkbox <input> or <select> drop-down and instead use your own entirely custom component solely for aesthetic reasons then you are producing worse software. There's nothing wrong with system-fonts, but your site will look ugly after your third-party font-host CDN shuts-down and turns into a walking CSRF factory.
> thoughtful typography with unexpected font pairings
The above fragment set my alarm-bells off. Yikes.
> scroll-triggered interactions
Not every web-page should be an Apple.com product brochure page. This is also a fantastic way to make your webpage horribly inaccessible.
------
The SKILL.md itself[3] grinds my gears too:
> Approach this as the design lead at a small studio known for giving every client a visual identity that could not be mistaken for anyone else's.
Claude has no way of knowing what designs are actually unique or not...
> For web designs, the hero is a thesis. Open with the most characteristic thing in the subject's world, in whatever form makes sense for it: a headline, an image, an animation, a live demo, an interactive moment
...this is exactly what everyone else's web-pages look like!
> For calibration: AI-generated design right now clusters around three looks: (1) a warm cream background (near #F4F1EA) with a high-contrast serif display and a terracotta accent; (2) a near-black background with a single bright acid-green or vermilion accent; (3) a broadsheet-style layout with hairline rules, zero border-radius, and dense newspaper-like columns
...I called this out weeks ago[4], lol.
and I could go on. This is all quite painful to read.
I would love more vision benchmarks! Once I asked the model to inspect a completely black picture and it hallucinated a nice wooden kitchen wall. Took me some time to figure out where the kitchen came from...
One of my friends (and BIL) own an architecture firm. They use AI to generate and quickly update renderings but they run into the equivalent of the 6 fingered hand problem. I sent him this article I wonder if the updated models can catch and fix mistakes made by previous models.
I recently used it at grocery stores in a foreign country. Photographed the whole aisle and told it to find Y (detergent, softener, glue, sour cream, whatever), at the same time recommend the best Y for whatever reason. Worked marvelously, including the cases where the object wasn't present and it told me there was nothing useful.
I asked then, can you crop the exact image of how does the item look like and where is it in the aisle - did that perfectly as well.
I will add that all frontier models were fine with such tasks from the early 2024's.
I've decided it's "good enough" after I saw it properly quote a string of text that was very roughly highlighted within a nested visual context. It also identified the context correctly (modal inside webapp inside screenshot of user desktop).
I understand why you would like to use an LLM for vision. I do it myself often enough. I don't understand however, why the pill detection and counting is included in this benchmark. That is a task which you would perform with OpenCV right?
In my personal mini benchmark minicpm-v-4.6 scores amazingly well. Its a 0.8B model which runs fine on many consumer hardware.
Generating datasets to train more efficient models is a common use case for VLMs, especially frontier ones. It makes it much cheaper to create that initial dataset and you can abuse the nondeterminism of LLMs to identify data for human review (if they don’t converge, escalate to a human).
In my practice Gemini models are far better than anything on the market in terms of vision, also it's worth to mention that current Gemini flash is 3.7, so it got 2 updates since 3.5 which beat GPT-5.6 Sol in this comparison.
If you're doing any kind of inference that is multi-modal and non-factual, opinions and biases will affect any kind of assessment of a visual that you provide to a model.
For example, a UI / UX professional being asked to appraise a website screenshot may determine that the image in question has "desirable" traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be "fashionable" with current UI trends.
As a design system engineer I usually have to fight against the taste of the designers. (And I consider it natural.)
But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.
Of course only if the design is achievable in the design system.
For 2 weeks I've been trying to get Codex to "outpaint" a wonderful image it generated as placeholder art for a level background.
After I increased the game's resolution, I asked it to increase the image's size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.
An average pixel-artist could have recreated the image and more within 2-3 days.
I agree. It did very well on an extremely challenging task.
I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass.
In addition, the poster itself also happened to contain similar clothing.
Ironically, the pill counting example selected to showcase "the best vision model" can be easily solved with OpenCV template matching, a technology created 25 years ago.
Anecdotal, opinion:
Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.
Assessing the subjective quality of a thing is in my experience one of the worst ways to use any LLM.
anthropic frontend-design skill does a great job with it.
Have you actually read the frontend design skill? It’s placebo at best. Very short and barely focused on design: https://github.com/anthropics/skills/blob/main/skills/fronte...
What an annoying time for GitHub to go down.
My exposure to Claude-produced UIs is limited, but I have started to notice certain design trends they tend to have in-common, which might be becoming hallmarks of AI-produced UIs - the same way we've started noticing the clichés of low-effort LLM-generated text.
FWIW, the summary-description[1] of "frontend-design"[2] gives me a few things to pick at:
> create polished code
Methinks only if you're using it with a very popular framework like React. What happens if you ask Claude to make the UI in WinForms or MFC?
> high-impact animations
That's bad UX 101 right there: animations in a UI exist as an affordance to the user, and never for its own sake (e.g. macOS's "genie" animation when you minimize a window to the dock exists so the user knows where they can restore the window from). The only people who actually want "high impact animations" in software are salespeople who want something for demo purposes.
> generic system fonts, predictable purple gradients, and cookie-cutter components.
This screams wanting to be different for the sake of standing-out, not because it results in a better software product; users benefit when their software fits-in with platform conventions: if you refuse to use a stock checkbox <input> or <select> drop-down and instead use your own entirely custom component solely for aesthetic reasons then you are producing worse software. There's nothing wrong with system-fonts, but your site will look ugly after your third-party font-host CDN shuts-down and turns into a walking CSRF factory.
> thoughtful typography with unexpected font pairings
The above fragment set my alarm-bells off. Yikes.
> scroll-triggered interactions
Not every web-page should be an Apple.com product brochure page. This is also a fantastic way to make your webpage horribly inaccessible.
------
The SKILL.md itself[3] grinds my gears too:
> Approach this as the design lead at a small studio known for giving every client a visual identity that could not be mistaken for anyone else's.
Claude has no way of knowing what designs are actually unique or not...
> For web designs, the hero is a thesis. Open with the most characteristic thing in the subject's world, in whatever form makes sense for it: a headline, an image, an animation, a live demo, an interactive moment
...this is exactly what everyone else's web-pages look like!
> For calibration: AI-generated design right now clusters around three looks: (1) a warm cream background (near #F4F1EA) with a high-contrast serif display and a terracotta accent; (2) a near-black background with a single bright acid-green or vermilion accent; (3) a broadsheet-style layout with hairline rules, zero border-radius, and dense newspaper-like columns
...I called this out weeks ago[4], lol.
and I could go on. This is all quite painful to read.
------
[1] https://claude.com/plugins/frontend-design
[2] https://github.com/anthropics/claude-plugins-official/tree/m...
[3] https://github.com/anthropics/claude-plugins-official/blob/2...
[4] https://news.ycombinator.com/item?id=49187385
What is a "non-normative UI block"?
Segments of the UI that don't conform to any other existing established design or conventions
areas that look weird
Penny sample shown looks like failed EXIF orientation registered by the model/harness. The coins are correctly marked, it's rotated 90 degrees.
I would love more vision benchmarks! Once I asked the model to inspect a completely black picture and it hallucinated a nice wooden kitchen wall. Took me some time to figure out where the kitchen came from...
I usually go to https://arena.ai/leaderboard/vision/pareto for a nice overview of current models.
In the third vision bench result, Sol is 100% correct but the expected has 1 error. Seems like an oversight.
In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.
One of my friends (and BIL) own an architecture firm. They use AI to generate and quickly update renderings but they run into the equivalent of the 6 fingered hand problem. I sent him this article I wonder if the updated models can catch and fix mistakes made by previous models.
All of your use cases are very advanced.
I recently used it at grocery stores in a foreign country. Photographed the whole aisle and told it to find Y (detergent, softener, glue, sour cream, whatever), at the same time recommend the best Y for whatever reason. Worked marvelously, including the cases where the object wasn't present and it told me there was nothing useful.
I asked then, can you crop the exact image of how does the item look like and where is it in the aisle - did that perfectly as well.
I will add that all frontier models were fine with such tasks from the early 2024's.
I've decided it's "good enough" after I saw it properly quote a string of text that was very roughly highlighted within a nested visual context. It also identified the context correctly (modal inside webapp inside screenshot of user desktop).
I understand why you would like to use an LLM for vision. I do it myself often enough. I don't understand however, why the pill detection and counting is included in this benchmark. That is a task which you would perform with OpenCV right?
In my personal mini benchmark minicpm-v-4.6 scores amazingly well. Its a 0.8B model which runs fine on many consumer hardware.
Generating datasets to train more efficient models is a common use case for VLMs, especially frontier ones. It makes it much cheaper to create that initial dataset and you can abuse the nondeterminism of LLMs to identify data for human review (if they don’t converge, escalate to a human).
I didn't expect Gemini 3.5 Flash to top basically every metric in this article.
In my practice Gemini models are far better than anything on the market in terms of vision, also it's worth to mention that current Gemini flash is 3.7, so it got 2 updates since 3.5 which beat GPT-5.6 Sol in this comparison.
Same. I scrolled back up to see if I read the title correctly. It's important to note that it is the best... OpenAI released. Not the best overall.
Does any popular NVR make a good use of LLMs (especially local models) getting decent at vision?
"Best iPhone ever" vibes.
My anecdotal evidence says its still as blind as any other model, it has no taste, no attention to any sort of detail.
How can a vision model have taste?
If you're doing any kind of inference that is multi-modal and non-factual, opinions and biases will affect any kind of assessment of a visual that you provide to a model.
For example, a UI / UX professional being asked to appraise a website screenshot may determine that the image in question has "desirable" traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be "fashionable" with current UI trends.
> if the interface elements have strong information hierarchy
...but that's an example of a UX/usability matter that can be assessed objectively and non-subjectively.
Replace taste with consistent if that helps you. Can it follow a design system...
As a design system engineer I usually have to fight against the taste of the designers. (And I consider it natural.)
But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.
Of course only if the design is achievable in the design system.
This is not my experience at all.
So, formulaic output…the opposite of taste
Not really. Compliance with the letter of the law doesn't mean the intent is complied with.
Still not quite as good as gemini.
Where are the Qwen benchmarks in this? I would be more interesting to see how Qwen performs.
For 2 weeks I've been trying to get Codex to "outpaint" a wonderful image it generated as placeholder art for a level background.
After I increased the game's resolution, I asked it to increase the image's size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.
An average pixel-artist could have recreated the image and more within 2-3 days.
I'm unsure why you're using an LLM to generate images. Don't we already have models (some made by the same company) that do this?
> it constantly keeps getting something wrong no matter what I tell it
This 100%
did you try segmenting it first?
I agree. It did very well on an extremely challenging task.
I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass.
In addition, the poster itself also happened to contain similar clothing.
You can see the reference images and its output in my writeup here: https://medium.com/@rviragh/gpt-5-6-sol-very-good-image-reco...
While a human can focus on the reflection easily, this is an enormous challenge for a vision model. It's very impressive.