Karpathy’s Pelican

260 points192 comments17 hours ago
YmiYugy

I don't think it's a bad way to benchmark new models, I just find it concerning that the author implies that "pelican on a bicycle" has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.

show comments
Lerc

I think I could tolerate 50 Shades of grey rendered in this style.

jmugan

A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)

show comments
bredren

I worked with an LLM to build a ~3D animation of the Back to the Future delorean Time Machine as a way to spice up the hero on a docs page.

That took a fair amount of custom tuning and I had to create a tuning view to get some of the behaviors right.

But it was enough fun that I generalized it to take in ~any scene description from a film. It goes out and gets more detailed descriptions and film stills if available but also takes custom stills if you provide them.

My test scene was the Gauntlet scene from Apocalypto. It is low fidelity but does a pretty amazing sequence with somewhat believable physics of the javelins etc.

Here is the docs page with the vertical takeoff / 88 miles an hour time travel: https://contextify.sh/docs

I can share some of the Apocalypto bit if anyone is interested.

show comments
djhworld

It would be interesting to see the models work on a book it hasn't been trained on yet. I guess sadly that means any book released very recently.

Definitely impressive demo, I do wonder though if the countless artwork, films, images etc produced over many decades around Lord of the Rings somewhat influenced the outcome of this though.

HarHarVeryFunny

It seems pretty clear that Anthropic models have been specifically trained to be good at generating three.js (JavaScript 3-D Graphics) code, so given current state of AI code generation in general, I don't find three.js models/animations as indicative of anything other than the model's ability to write three.js code.

When Fable was first released the day-1 demos of it on Twitter (presumably from people who were given early access, and/or Anthropic employees) were pretty much 100% three.js stuff. Yes, it looks nice, but it doesn't tell me any better than an Erdos proof whether the LLM will be able to run my vending machine.

show comments
try-working

this is not a good benchmark for models, but it's great if you're optimizing for attention on twitter because video content and 3d animations perform best on social media.

a real benchmark is instead running evals on your own traces, and building a cost/quality/speed profile for models based on real workloads. but it doesn't get you a shiny video you can post on twitter.

jcims

I’d like to see a human one shot a pelican on a bicycle in raw svg.

show comments
qwertox

I'd rather have them battle on the topic "Who builds a better Google Wave for LLM chats" to explore the space of how AI studios could be.

"Getting Started with Google Wave": https://www.youtube.com/watch?v=eKUAqNGVwX0

show comments
dundarious

I can forgive the modeling being godawful jank (windows floating in the air, disconnected from the house). But I expected it to have a better understanding of the text. Instead, we have Bilbo's "disappearance" interpreted as him magically transporting or cloaking, and similarly for his reappearance.

show comments
toolslive

Reading the title, I was thinking "Karpathy? I don't know this chess player." (The Pelikan is a well known chess opening, and famous chess players often have book titles like "X's Y" where X is the player, and Y is the opening)

trentor

I always thought of the Pelican more of like a gimmicky quick test. There are people who took it as a serious benchmark for overall model performance?

show comments
swe_dima

In my experience SVGs are still too hard for LLMs.

I gave Fable a jpeg and asked to draw an SVG, using a loop that renders the SVG into an image so Fable can inspect it.

Results looked like drawing of a 5 year old.

informal007

One difference for human to understand the video is that we only care the changes on a picture compare to LLM

Waterluvian

Speaking of benchmarks has anyone given AIs Where’s Waldo pages and asked it to find Waldo?

I’ve been trying it on them all and can’t find one that does it consistently. The best will tell me they can’t. The worst confidently point out one of countless Waldo-likes.

baron816

IMO, the area where AI is going to be most useful over the next couple years is in developing manufacturing processes top to bottom. Maybe a million token budget is too small, but something like "design me a sneaker and all the equipment to manufacture it autonomously".

show comments
hooloovoo_zoo

I suspect LotR is a singularly unrepresentative choice here considering how much info exists about it.

siliconc0w

There is a tipping point between procedurally generating everything in SVG to maybe giving them tool access to something like 3dsmax (or having them build and then use a tool to do the thing vs doing the thing).

show comments
fzeindl

Regarding the argument about LLMs having difficulties auditing their work:

I wonder whether we are entering the era of throwaway software. Just like cheap plastics and improved processes has enabled us to rapidly manufacture anything we want for a very low price, maybe LLMs give us the same for software. Produce it cheaply and if it breaks throws it away and reproduce it.

show comments
xyzsparetimexyz

How much are the hobbit houses described in the book? The ones here look exactly like the movie

show comments
matsemann

I'm pretty tired of the "Y made this game in Z tokens" all over the internet last week. They look impressive, and it's cool that it's even possible, but they're useless as games. None of them are any fun. They're like the most boring variant of basic controllers you can imagine. None have any cool mechanics. None have any tweaks made from hours and hours of testing. All have the same cel-shader.

show comments
eichin

Is anyone else getting "mongodb is webscale" vibes? (Except 16 years ago that was a lot smoother, because it used some sort of "render this conversation" engine...)

sinaatalay

On consumer devices, AI communicates with us through speakers and screens. Screens are the richer medium, so most consumer AI innovation will happen there.

Computer graphics will have enormous applications because they are directly controllable by LLM-generated code. Video models are probabilistic and less suitable when precision matters. In education, for example, we need exact visuals. If an AI wants to plot y = sin(x), it should generate the precise graph through computer graphics rather than approximate it with a video model.

show comments
cocoa19

We must not be using the same opus 5, because if I tried to generate this it would refuse based on copyright grounds.

show comments
dekhn

I'd like to see the Silmarillion, specfically both Ainulindalë and the Fall of Numenor. At this point a visual model would probably produce something better than Amazon (but presumably not Jackson).

informal007

it shows the possibility that SVG replace PNG/JPG even video.

fwlr

I really dislike this AI programming thing of “Mr LLM, go slam your face into the problem until there’s no problem left, then call me back”. (Not sure if it’s a recent trend or a fundamental nature.)

It always brings to my mind some words from Rich Hickey:

    I think we’re in this world I’d like to call “guardrail programming”. It’s really sad: we’re like, “I can make change because I have tests!”. Who does that? Who drives their car around, banging against the guardrails, saying “whoah, I’m so glad I have these guardrails so I can make it to the show on time!”
I don’t think I really have a point to make here, other than it just feels like someone’s released a bunch of carnival bumper cars onto the highways.
show comments
barrenko

This has started to feel a bit like the beginning of railroads and then the steampunk fiction of "let's just build railroads to everywhere". We don't need it and there's no use for it.

As with painting, after a while there's nothing really new to paint, we genuinely need 0 new software. We need to fix our broken physical world, our social lives, our kids and what's left of our democracies.

This software crap is done, leave it to the nerds.

show comments
mold_aid

The tilde thing remains uniquely obnoxious in a field that seems want to mangle language for fun, so that's innovative I guess

skybrian

Still images seem like a better quick test because we can see them at a glance. Maybe ask it to make a comic?

hkalbasi

This makes me think about using a game engine and a coding agent instead of current video generation AIs. It will probably cost much more, but it will have almost zero consistency problems. Is this line explored?

OtherShrezzing

> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom

This is an odd take, given that Karpathy is certainly aware that the LotR films absolutely did create Bag End in digital format; that their creation was outstandingly high quality; and that Claude’s output here very obviously “leans heavily” on their prior art.

wiradikusuma

Do you guys notice that LLM can create fancy viz/animations by coding them instead of leveraging what we humans usually use (e.g Lottie, After Effects)?

I wonder if Flash is still popular... LLM can use that instead...?

dofm

Anthropic spokesman [0] Andrej Karpathy is here to tell you about token-wasting loops, and insists on the weird idea that they are "~free", when in fact, they are fuelled by expensively burning investor money.

[0] Seriously. Get used to mentally prefixing his and Boris Cherny's name like this, every time you see them quoted. These people are speaking while employed; there is no chance they are not aligned with the employers who will make them wealthy. The tech industry does like to pretend that for some reason AI people, uniquely, speak thoughts unbiased and for themselves or even for science or humanity.

serf

you don't really need screenshots if you have an engine expressive enough for the scene generation while ensuring the visual appearance of the engine output itself is feasible.

that's why these things are actually pretty good at openscad/freecad/F360 mcps , the visual reality is enforced and guaranteed by rigor in the interpretation engine that is anchored to human physical reality.

croes

> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom

There are people in their right mind who would do that and their are already examples of people who did similar things.

But maybe not in the future if people would confuse all the effort with AI

xnx

AI is now somewhere between "Money for Nothing" (https://www.youtube.com/watch?v=wTP2RUD_cL0) and "Knick Knack" (https://www.youtube.com/watch?v=9uhM_SUhdaw) in capabilities.

throwaway89864

It may make sense to switch this to USD/Omniverse.

mikojan

After watching this video I am absolutely positive that the issue is not a lack of stamina in humans. It is that humans have the capacity to realize that this is a bad idea long before they complete it.

xg15

> I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it.

I think it's interesting that the "Bag's End" interpretation in the video clearly looks like the one from the movies, but generated here as a three.js 3D asset.

It makes sense that the movies (or shots/frames from them) were in the training data, and I can also easily imagine an association in concept space between the textual description of Bag's End and the frames from the movie.

But how on earth does the model then go on and convert the latent representation of those images into coordinates for a 3D mesh, without ever even restoring the image? In what kind of representation are the images from the movies stored that it can do that?

wslh

If you like this check: https://news.ycombinator.com/item?id=47400868 it can be used to generate animations (not games) as well.

show comments
Gooblebrai

I can't believe the video demo is $10

quantumleaper

I'm sad that Andrej Karpathy went from being one of the most reasonable, trusted, and credible voices in AI to peddling marketing slop for Anthropic.

8 months ago, he was (very reasonably) claiming that reliable agents are at least a decade away, but this now goes against the interest of his employer, so the narrative has been changed.

show comments
bbstats

This is awful

blitzar

I think the pelican test is better.

show comments
andy99

Benchmarks like the pelican thing are about correlation with “how good the model is”. Better models produce better pelicans.

It’s a useful benchmark (aside from being “cute”) because of its simplicity, both in how many output tokens it takes (though I understand some models think a lot now to do it) and how easily one can subjectively judge. It’s this efficient as a benchmark of performance.

Making a long video takes way more tokens, and presumably is a lot tougher to easily compare. swillison has a presentation that’s pelicans from 2023-present (roughly) showing the progression. Imagine “lord of the rings videos from 2026-2029” or whatever, it would take a long time to watch and be harder to judge, and probably just end up being a comparison of screenshots anyway.

TLDR I feel like the post misunderstands the role of the pelican thing though if find it very hard to believe he really doesn’t understand, so maybe I’m missing something.

stackedinserter

It would be better to ask model to render segmented 3d, with placeholders, like magenta is water, blue is sky, green is grass, purple is Frodo's face, etc, then pass the result through img2img model to properly "render" it.

c0rruptbytes

perfect benchmark to burn more tokens - convenient

epolanski

I wish there was a timeline where I never ever had to see the pelican SVG test ever again.

forrestthewoods

As a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment.

There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.

In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.

My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.

Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.

nozzlegear

Don't miss Elon Musk's reply:

> @elonmusk 13h

> Yah

> 158 replies, 74 reposts, 1400 likes

Thanks Elon, you goofy fuck

https://xcancel.com/elonmusk/status/2083761408932458568

show comments
shapefrog

"Check the current situation and make a new Iran Lego (tm) truth bomb video."

miltonlost

Tech bros continue wasting money to make the absolute worst art

show comments
andrewstuart

This is equally bad as a pelican test.

LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.

So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).

You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.

Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.

No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.

Any test that anyone comes up with must clearly state the context and how the outcome is measured.

hn22fazjsv

Screenshotting for later

matchagaucho

It's difficult to think in exponentials.

But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.

show comments