I'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.
I’ll never not be amazed that we can type some words and get those results back out.
However I do agree that the results are much worse when you try to use them. They look great in screenshots and video clips which makes them perfect for content farmers.
All of the LLM generated games I’ve played have been really bad to play, though. I even tried my hand at a simple game, thinking I could iterate on it with prompts to fix some rough edges. After the initial productivity burst every change turned into a slog of tokens with one thing changing and something else breaking it. I would try to use my remaining weekly token budget across Anthropic and OpenAI to refine it at the end of every week but after a couple weeks it felt like I wouldn’t be getting anywhere without scrapping it and going back to having the LLM build it one step at a time with my careful instruction.
Which, in retrospect, is the only way I can get usable output of an LLM for anything complicated, so it’s not surprising. It’s a fun reality check project though.
LLMs are bad at creative work and I don’t see them improving any time soon. Try asking an agent to write a story about raccoons. It will almost certainly involve either stealing food or raiding trash with a 50% chance of having a character named Pip. If a location is mentioned, it’ll be Elm Street.
I guess this is the average story and, similarly, the average game is boring and predictable
Seriously, what is the deal with Pip! I was experimenting with childrens' fiction more than a year ago and 9 times out of 10 you'd get a character named Pip unless you were specific in the prompt not to do so.
Unfortunately people are saying they are fun for the purposes of ragebait that goes viral, which is half the reason people intentionally post AI slop on social media.
LLMs currently have 0 imagination, I have strong doubts that will be solved any time soon even if they continue to become superhuman at everything else. They just have the creative instincts of a 50 year old accountant.
Unfortunately the Phaser framework has decided to go all-in on this route, and it's really disappointing.
These one-shot products aren't games. They're barely even demos. I don't even know what to call them. For a mature framework like Phaser to sell-out like this and create a vibecoded platform for vibecoded games is shocking.
It’s called “demo porn” and it’s really obnoxious and borderline offensive to people that take the craft and art form of video games seriously.
If I see one more “one shot MMO” where you just walk around and do absolutey nothing or another menu slop idle battler or rogulike deck builder I’m going to go Postal in Minecraft.
Anthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money.
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or worse, for science or humanity.
I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.)
It always brings to my mind some words from Rich Hickey:
I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!”
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.
I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.
A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.
I used it to plan a group beach trip back when it came out. We had a single page shared with everyone going on the trip. The page had live shopping lists, maps, weather forecasts, and other snippets of useful information. Now in 2026 I still can't think of any single technology that provides the same utility. Although to be fair, this might be a case of rosy retrospection.
I'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).
Regarding the argument about LLMs having difficulties auditing their work:
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
I'm sure this will become the standard. And plastic is an excellent analogy. Maybe we can take it a step further and compare it to on-demand 3D printing.
Why would anyone still use off-the-shelf software when they can have a system that has access to all data, can transform it into any form, and can export it in any format?
After years of thinking that I needed to develop a decent movie management system for my own films or a columnar browser for large CSV files, Claude and Qwen each delivered exactly what I needed in just a day.
Does this make sense with economics of software though? Throwaway products compete with more durable versions of the same product because there is a cost per unit produced that can be minimized by using cheaper materials or production processes that cut corners. With software there is no cost per unit. There might be a market for one-off software that serves a very specific purpose where throwaway software can compete with adapting more carefully engineered software to that purpose, but I am not convinced there is a lot value in this market.
It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
This makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?
IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
I think embodied AI or autonomous experimentation will need to make a lot more progress before that kind of thing is possible.
Asking AI to design real world objects doesn't work very well because all of its tests involve proxies and thus miss things that are glaringly obvious when the object is actually built.
Yes but the second order effect of this is that the cost of the tooling goes up since it is now the bottleneck, and therefore the shoemakers that survive do it off of technical complexity, branding, and regulatory capture.
> "design me a sneaker and all the equipment to manufacture it autonomously"
We already have sneaker designs and the equipment to manufacture them. Whatever it spits out is going to be, at best, a mediocre clone of something that already exists. What exactly is the point?
> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it.
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
you don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
Do you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)?
I wonder if Flash is still popular... LLM can use that instead...?
I'd like to see tests of things the current AIs are bad at, like drive a car. Or maybe take instruction to complete some novel activity to test how well they can learn.
As a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment.
There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.
In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.
My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.
Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
So the LLMs do know the first paragraph for any novel! Probably verbatim, as we saw in the 2023 models. That verbatim giveaway has been beaten out of them to keep up appearances.
Theft of what? Why - aside from the fact that your accusation are irrelevant to the submission and arguments absent -, you do not know verbatim paragraphs (we do)?
You cannot come and place your personal positions as assumptions. To me, there is absolutely no theft. And we cannot play a game of "Yes!"//"No!" here.
Objections are arguments. That post did not start an argument - it was as ideological as the Brigades. Devoid of what we want to have here (I believe).
I have stated and do state: what is published is assumed as read (only, probably not read for lack of resources). If it is in the libraries, it is there to be read.
I'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic.
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
> 8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
Or maybe over these 8 months agents improved a lot? You know, few years ago many AI experts predicted that things we are routinely doing now with AI are decades away. I mean how can you look at this post and not be impressed? It's insane what AI is currently capable of.
He makes a very good point here about llms or lrms not being good at taking video as inputs, and i think he's signaling his intent to help change that. He's pointing at an open problem. This isn't a carefully structured blog post, it's just a tweet.
What narrative has changed ? I just see an experiment with a very perfectible result and some reflections about what capacity are currently missing to have better results.
But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
I don't really get the desire for hyper-personalized entertainment. People are very good at pointing out things they dislike, but not very good at coming up with how to fix them (common wisdom in game design).
On top of that, a decent chunk of the joy of entertainment is the social aspect.
If that becomes a thing, that will be fun 2 weeks, then become a gimmick only a small niche will be using. What’s the point in having a hyper personalized entertainment? People want to experience the games, movies, tv shows, books others created and are also experiencing
I very much doubt that. We now have nearly-perfect AI image generation and it hasn't really changed the nature of human expression. I don't see my friends getting wildly creative. I mostly see it used for spammy blogs, spammy books, and cringeworthy corporate marketing - basically, a negative signal, rather than "ooh, AI image, I'm in for a treat". Is your experience different? If not, what changes with moving images?
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want, and they don't want to be in charge of content production (even if you have an LLM do most of the work). I don't doubt that there are niches for it, especially stuff like porn, and I'm sure that pros (game studios, film studios) will leverage AI more and more, but I suspect that most of us will just want to sit on the couch, watch Spiderman XVIII, and then be able to talk about that shared Spiderman XVIII experience with all our friends.
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want
1:1 AI entertainment probably won't be a cold start experience.
Much like the "Choose Your Own Adventure" books of the 80's, consumers choose a baseline template, customize the characters, and interact at various points within the plot.
I'm mostly seeing the impact of GenAI images in everydays life for example in the small posters small associations or individuals usually attach on streets to promote some small local event. We went from just text created with PowerPoint and maybe some stock image just a Google search away to now images depicting the topic closely. But yeah, it's not like a revolution.
And this personal media thing, yeah maybe for terminally online persons that are REALLY into a sub-genre but otherwise, it's too much effort, I agree. Until we get machines that can read our (subconscious) mind, that will not exist.
Not even close. Current AIs have very poor spatial awareness, they can generate some kind of scene but they can't tell you where objects are in the scene, nor can they move objects into different places. They're very useful but it's difficult to be very creative with them because they can't update an image to make it more aligned with your vision for what should be in the image.
Amazon recently kiboshed a new Stargate television series. The fan community was quite heartbroken because the Amazon had involved some of the writers from the original television series and apparently one of the reasons that they cancelled the show before it made it to production was that the felt that the proposed story would appeal too much to the fans and not new people.
The fans were real heart broken about this but I think you're right on where this is going. We're not going to be seeing the dominance of centrally produced content like this for much longer, like sure, I think there will be big blockbusters will stick around, but I think the day is coming where media becomes a choose your own adventure sort of scenario.
It'll be interesting to see where this scales to. there will definitely be some amazing solo projects but we'll also see the like 4 player co-op version of productions and then the larger mine-craft server 'Minas Tirith' scale ambitious projects that involve a few dozen people. And of course passive consumers will remain a thing, or people who just provide some suggestions or nudges for what they'd like to see others make.
But I don't think it'll be dominated by big companies like Disney, Netflix, Amazon or Paramount.
I'm sure a lot of them will suck but it'll be neat to see the inevitable Seinfield - Star Trek Voyager cross over episodes.
Elaine and B'Elanna Torres feud after a transporter accident leaves the crew stranded the delta quadrant. Jerry attempts to date 7of9 but is rebuffed as she finds Kramer's quirky bluntness more relatable. George panics after someone compares him to Neelix.
I'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.
I’ll never not be amazed that we can type some words and get those results back out.
However I do agree that the results are much worse when you try to use them. They look great in screenshots and video clips which makes them perfect for content farmers.
All of the LLM generated games I’ve played have been really bad to play, though. I even tried my hand at a simple game, thinking I could iterate on it with prompts to fix some rough edges. After the initial productivity burst every change turned into a slog of tokens with one thing changing and something else breaking it. I would try to use my remaining weekly token budget across Anthropic and OpenAI to refine it at the end of every week but after a couple weeks it felt like I wouldn’t be getting anywhere without scrapping it and going back to having the LLM build it one step at a time with my careful instruction.
Which, in retrospect, is the only way I can get usable output of an LLM for anything complicated, so it’s not surprising. It’s a fun reality check project though.
LLMs are bad at creative work and I don’t see them improving any time soon. Try asking an agent to write a story about raccoons. It will almost certainly involve either stealing food or raiding trash with a 50% chance of having a character named Pip. If a location is mentioned, it’ll be Elm Street.
I guess this is the average story and, similarly, the average game is boring and predictable
Ironically, they're getting worse at creative work because all the RL is collapsing their distributions.
Seriously, what is the deal with Pip! I was experimenting with childrens' fiction more than a year ago and 9 times out of 10 you'd get a character named Pip unless you were specific in the prompt not to do so.
I recently saw a study on this actually. https://arxiv.org/abs/2605.26492
I think the trick is to use LLMs like scalpels instead of hammers. That said, you can place a scalpel in the hands of a deterministic (or not) robot.
Elias Thorne!
We're all just impressed that it's even possible. No one is saying these are fun.
I guess the next question is if they can be made fun without too much additional work with a human guiding the AI.
How well encoded is "good game feel" in the weights of an LLM. I am guessing not very well.
> the next question is if they can be made fun without too much additional work
They can't. If you think about how these things are trained it's blatantly obvious fun is an impossible metric to optimize them for
Training impossibilities aside, how would you even make an optimization loop for "fun"?
Brain-computer interface, maybe. Hook a million play-testers up to the output and have it iterate.
..feels like there was a black mirror episode about that though
Unfortunately people are saying they are fun for the purposes of ragebait that goes viral, which is half the reason people intentionally post AI slop on social media.
LLMs currently have 0 imagination, I have strong doubts that will be solved any time soon even if they continue to become superhuman at everything else. They just have the creative instincts of a 50 year old accountant.
> They just have the creative instincts of a 50 year old accountant
That's oddly specific and it'd really hurt my cousins feelings :)
Unfortunately the Phaser framework has decided to go all-in on this route, and it's really disappointing.
These one-shot products aren't games. They're barely even demos. I don't even know what to call them. For a mature framework like Phaser to sell-out like this and create a vibecoded platform for vibecoded games is shocking.
It’s called “demo porn” and it’s really obnoxious and borderline offensive to people that take the craft and art form of video games seriously.
If I see one more “one shot MMO” where you just walk around and do absolutey nothing or another menu slop idle battler or rogulike deck builder I’m going to go Postal in Minecraft.
It's the flashy game/movie trailer equivalent for LLMs. Completely unrelated to the real experience, but good for marketing.
It's blockchain for gaming all over again, putting the cart before the horse.
you're tired because you lack imagination.
It’s the people spamming these shitty games that lack imagination.
Anthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money.
[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or worse, for science or humanity.
I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.)
It always brings to my mind some words from Rich Hickey:
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.> I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”.
Not difficult to see why the employees of AI firms are thrilled with it though, eh?
I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.
That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.
But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.
My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.
Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs
I can share some of the Apocalypto bit if anyone is interested.
I am interested, I'd love to see! I'm waiting for the day my dad's self published books become self produced movies!
I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.
A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.
I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.
"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0
i remember this failing, but this product looks pretty useful in 2026
I used it to plan a group beach trip back when it came out. We had a single page shared with everyone going on the trip. The page had live shopping lists, maps, weather forecasts, and other snippets of useful information. Now in 2026 I still can't think of any single technology that provides the same utility. Although to be fair, this might be a case of rosy retrospection.
Notion pages are pretty good for shared trip planning docs. Although maybe without so many live updating widgets.
It looked pretty useful in 2009. Failing to push it was baffling then too.
It was open sourced as Apache Wave when Google shut it down.
Years later Apache moved it to read only because of low community activity.
The archived git repo on GitHub remains available to clone and revive as a fork.
https://github.com/apache/incubator-retired-wave
This doesn't explain the lack of marketing. Apache isn't exactly known for pushing tech.
It was awesome when released. I used it a lot, the multiplayer experience was awesome, and the mix of document-forum-wiki is still something I miss
notion isn't too far off
wave failed for weird google organizational reasons far more than anything inherent to the product or tech
I'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).
Regarding the argument about LLMs having difficulties auditing their work:
I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.
I'm sure this will become the standard. And plastic is an excellent analogy. Maybe we can take it a step further and compare it to on-demand 3D printing.
Why would anyone still use off-the-shelf software when they can have a system that has access to all data, can transform it into any form, and can export it in any format?
After years of thinking that I needed to develop a decent movie management system for my own films or a columnar browser for large CSV files, Claude and Qwen each delivered exactly what I needed in just a day.
Yes. It's fantastic for one-off tasks.
Does this make sense with economics of software though? Throwaway products compete with more durable versions of the same product because there is a cost per unit produced that can be minimized by using cheaper materials or production processes that cut corners. With software there is no cost per unit. There might be a market for one-off software that serves a very specific purpose where throwaway software can compete with adapting more carefully engineered software to that purpose, but I am not convinced there is a lot value in this market.
We must not be using the same opus 5, because if I tried to generate this it would refuse based on copyright grounds.
Karpathy works for Anthropic so he obviously has full access to everything, and doesn’t have to pay for token burn either.
It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.
When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.
This makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?
IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".
I think embodied AI or autonomous experimentation will need to make a lot more progress before that kind of thing is possible.
Asking AI to design real world objects doesn't work very well because all of its tests involve proxies and thus miss things that are glaringly obvious when the object is actually built.
You think the bottleneck in creating a sneaker factory is not having an LLM to tell you what machines to buy?
Yes but the second order effect of this is that the cost of the tooling goes up since it is now the bottleneck, and therefore the shoemakers that survive do it off of technical complexity, branding, and regulatory capture.
> "design me a sneaker and all the equipment to manufacture it autonomously"
We already have sneaker designs and the equipment to manufacture them. Whatever it spits out is going to be, at best, a mediocre clone of something that already exists. What exactly is the point?
Gotta love when a techbro just says some complete nonsense like this with total confidence
Grok build me a spaceship to mars, make no mistakes
I always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?
"Draw a pelican on a bicycle" is not a serious benchmark.
"Draw an animation of this long ass scene from a movie, and only call me when everything works e2e" can be.
Still images seem like a better quick test because we can see them at a glance. Maybe ask it to make a comic?
How much are the hobbit houses described in the book? The ones here look exactly like the movie
It may make sense to switch this to USD/Omniverse.
> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom
There are people in their right mind who would do that and their are already examples of people who did similar things.
But maybe not in the future if people would confuse all the effort with AI
> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it.
I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.
It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.
But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?
you don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible.
that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.
I can't believe the video demo is $10
Do you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)?
I wonder if Flash is still popular... LLM can use that instead...?
perfect benchmark to burn more tokens - convenient
This is awful
I think the pelican test is better.
I'd like to see tests of things the current AIs are bad at, like drive a car. Or maybe take instruction to complete some novel activity to test how well they can learn.
As a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment.
There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.
In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.
My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.
Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.
"Check the current situation and make a new Iran Lego (tm) truth bomb video."
Tech bros continue wasting money to make the absolute worst art
> wasting money
Testing, assessing, tasting...
We also do it when we build other things - this is just a different scale.
Do you think this was supposed to be art? Wow.
I wish there was a timeline where I never ever had to see the pelican SVG test ever again.
So the LLMs do know the first paragraph for any novel! Probably verbatim, as we saw in the 2023 models. That verbatim giveaway has been beaten out of them to keep up appearances.
Thanks for confirming:
https://web.archive.org/web/20260802165914/https://xcancel.c...
You can downvote like in all other ClosedAI threads, but justice will come for the thieves eventually.
Theft of what? Why - aside from the fact that your accusation are irrelevant to the submission and arguments absent -, you do not know verbatim paragraphs (we do)?
You cannot come and place your personal positions as assumptions. To me, there is absolutely no theft. And we cannot play a game of "Yes!"//"No!" here.
By the way: are we having a surge of this?
> Theft of what? > By the way: are we having a surge of this?
We do have a surge of pro-AI sealions, yes. Any objection is countered with one or more three word questions.
> pro-AI sealions, yes
Very devoid of intelligence note.
> Any objection
Objections are arguments. That post did not start an argument - it was as ideological as the Brigades. Devoid of what we want to have here (I believe).
I have stated and do state: what is published is assumed as read (only, probably not read for lack of resources). If it is in the libraries, it is there to be read.
Try reading what Karpathy said again before reinforcing your biases.
Yeah, pretend that you never used 2023 models that quoted everything verbatim. To put it in a way that you'll understand:
No bias. No speculation. Just facts!
You literally did not read the article.
I'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic.
8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
> 8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.
Or maybe over these 8 months agents improved a lot? You know, few years ago many AI experts predicted that things we are routinely doing now with AI are decades away. I mean how can you look at this post and not be impressed? It's insane what AI is currently capable of.
He makes a very good point here about llms or lrms not being good at taking video as inputs, and i think he's signaling his intent to help change that. He's pointing at an open problem. This isn't a carefully structured blog post, it's just a tweet.
What narrative has changed ? I just see an experiment with a very perfectible result and some reflections about what capacity are currently missing to have better results.
Don't miss Elon Musk's reply:
> @elonmusk 13h
> Yah
> 158 replies, 74 reposts, 1400 likes
Thanks Elon
https://xcancel.com/elonmusk/status/2083761408932458568
Can someone please translate that expression? What would that mean?
Affirmation?
Well, in that case - if it is just a "hear, hear" - I do not see why the utterance from that actor would be of note.
Maybe nozzlegear wanted to suggest some importance on Musk remaining a bet-ter on the general tech, regardless of the competition?
Concerning
It's difficult to think in exponentials.
But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.
I don't really get the desire for hyper-personalized entertainment. People are very good at pointing out things they dislike, but not very good at coming up with how to fix them (common wisdom in game design).
On top of that, a decent chunk of the joy of entertainment is the social aspect.
If that becomes a thing, that will be fun 2 weeks, then become a gimmick only a small niche will be using. What’s the point in having a hyper personalized entertainment? People want to experience the games, movies, tv shows, books others created and are also experiencing
I very much doubt that. We now have nearly-perfect AI image generation and it hasn't really changed the nature of human expression. I don't see my friends getting wildly creative. I mostly see it used for spammy blogs, spammy books, and cringeworthy corporate marketing - basically, a negative signal, rather than "ooh, AI image, I'm in for a treat". Is your experience different? If not, what changes with moving images?
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want, and they don't want to be in charge of content production (even if you have an LLM do most of the work). I don't doubt that there are niches for it, especially stuff like porn, and I'm sure that pros (game studios, film studios) will leverage AI more and more, but I suspect that most of us will just want to sit on the couch, watch Spiderman XVIII, and then be able to talk about that shared Spiderman XVIII experience with all our friends.
Ultimately, most people don't have ideas for the kinds of personalized entertainment they want
1:1 AI entertainment probably won't be a cold start experience.
Much like the "Choose Your Own Adventure" books of the 80's, consumers choose a baseline template, customize the characters, and interact at various points within the plot.
I'm mostly seeing the impact of GenAI images in everydays life for example in the small posters small associations or individuals usually attach on streets to promote some small local event. We went from just text created with PowerPoint and maybe some stock image just a Google search away to now images depicting the topic closely. But yeah, it's not like a revolution.
And this personal media thing, yeah maybe for terminally online persons that are REALLY into a sub-genre but otherwise, it's too much effort, I agree. Until we get machines that can read our (subconscious) mind, that will not exist.
> We now have nearly-perfect AI image generation
Not even close. Current AIs have very poor spatial awareness, they can generate some kind of scene but they can't tell you where objects are in the scene, nor can they move objects into different places. They're very useful but it's difficult to be very creative with them because they can't update an image to make it more aligned with your vision for what should be in the image.
Judging by the Seedance 2.5 demos today, I’d say it’s not that many orders of magnitude away now.
I kind of like a sense of community even if it means losing out on personalisation.
Amazon recently kiboshed a new Stargate television series. The fan community was quite heartbroken because the Amazon had involved some of the writers from the original television series and apparently one of the reasons that they cancelled the show before it made it to production was that the felt that the proposed story would appeal too much to the fans and not new people.
The fans were real heart broken about this but I think you're right on where this is going. We're not going to be seeing the dominance of centrally produced content like this for much longer, like sure, I think there will be big blockbusters will stick around, but I think the day is coming where media becomes a choose your own adventure sort of scenario.
It'll be interesting to see where this scales to. there will definitely be some amazing solo projects but we'll also see the like 4 player co-op version of productions and then the larger mine-craft server 'Minas Tirith' scale ambitious projects that involve a few dozen people. And of course passive consumers will remain a thing, or people who just provide some suggestions or nudges for what they'd like to see others make.
But I don't think it'll be dominated by big companies like Disney, Netflix, Amazon or Paramount.
I'm sure a lot of them will suck but it'll be neat to see the inevitable Seinfield - Star Trek Voyager cross over episodes.
Elaine and B'Elanna Torres feud after a transporter accident leaves the crew stranded the delta quadrant. Jerry attempts to date 7of9 but is rebuffed as she finds Kramer's quirky bluntness more relatable. George panics after someone compares him to Neelix.