I recently switched over to using primarily Bullet for my projects and the speed of it makes it very nice to get projects off the ground quickly and work properly. I also found that Adi and Alex make updates very regularly based on some of the feedback i've submitted to their feedback tab. Good work on this product guys! I'm excited to see how it develops in the future.
docheinestages
Your landing page is very hard to read. The font size is literally 10px for some content, while animations distract the reader.
Dear YC, please force the startups to dedicate some of the 500k funding for standard web design.
seizethecheese
This is a promising direction! Unfortunately, I think the benchmark result here is essentially meaningless.
I recently discovered this same lesson the hard way. I was trying to get a multi-agent system I was building to improve upon GPQA Diamond scores (system here: http://pellmell.ai). No matter how hard I tried, I could not get any lift. When Fable 5 dropped, it also did not improve upon Opus, and I realized my mistake. The benchmark was saturated!
Now, looking at the result here, I see a similar pattern. Fable is not better than Opus, and the score is ~95%. Notably, this post omits which subagent is being used. Why? An intellectually honest way to tell if this thing really works would be to run that agent and report its score and cost as well.
Going back to my GPQA Diamond lesson, you can see here how a saturated leaderboard behaves https://artificialanalysis.ai/evaluations/gpqa-diamond. Fable gets 92.6% for $0.22 per task while several models score higher for $0.01. I could easily publish a router that “enhances Fable on GPQA Diamond” showing improved score for lower cost, just by implementing a router that picks the model at random!
show comments
0kk33
Congrats on the launch.
I can't find which model providers are supported?
Are you calling the claude-code CLI directly and make bullet usable with anthropic subscriptions like orca or herdr?
I guess however this is a harness and it needs to connect to the API?
show comments
hmokiguess
> P.S: we hid a code on the website, see if you can unlock the secret page at the footer, all built with Bullet
Is it even hidden when your AI ends up tagging it with `aria-label="Hidden secret code"`? Lol. Fun mini game.
show comments
myshapeprotocol
Focusing on execution speed for coding agents is the right bottleneck to tackle. Exciting launch.
# Share chats with Bullet — helps us improve model routing and answer quality
Enabled by default.
show comments
tontinton
Are you freaking kidding me with YC throwing money at something like this? I guess I can fund raise just by having built https://maki.sh, and months ahead of other founders too...
I recently switched over to using primarily Bullet for my projects and the speed of it makes it very nice to get projects off the ground quickly and work properly. I also found that Adi and Alex make updates very regularly based on some of the feedback i've submitted to their feedback tab. Good work on this product guys! I'm excited to see how it develops in the future.
Your landing page is very hard to read. The font size is literally 10px for some content, while animations distract the reader.
Dear YC, please force the startups to dedicate some of the 500k funding for standard web design.
This is a promising direction! Unfortunately, I think the benchmark result here is essentially meaningless.
I recently discovered this same lesson the hard way. I was trying to get a multi-agent system I was building to improve upon GPQA Diamond scores (system here: http://pellmell.ai). No matter how hard I tried, I could not get any lift. When Fable 5 dropped, it also did not improve upon Opus, and I realized my mistake. The benchmark was saturated!
Now, looking at the result here, I see a similar pattern. Fable is not better than Opus, and the score is ~95%. Notably, this post omits which subagent is being used. Why? An intellectually honest way to tell if this thing really works would be to run that agent and report its score and cost as well.
Going back to my GPQA Diamond lesson, you can see here how a saturated leaderboard behaves https://artificialanalysis.ai/evaluations/gpqa-diamond. Fable gets 92.6% for $0.22 per task while several models score higher for $0.01. I could easily publish a router that “enhances Fable on GPQA Diamond” showing improved score for lower cost, just by implementing a router that picks the model at random!
Congrats on the launch. I can't find which model providers are supported? Are you calling the claude-code CLI directly and make bullet usable with anthropic subscriptions like orca or herdr?
I guess however this is a harness and it needs to connect to the API?
> P.S: we hid a code on the website, see if you can unlock the secret page at the footer, all built with Bullet
Is it even hidden when your AI ends up tagging it with `aria-label="Hidden secret code"`? Lol. Fun mini game.
Focusing on execution speed for coding agents is the right bottleneck to tackle. Exciting launch.
Skip signup:
Cmd+Option+I > Console > 'allow pasting'
const onboarding = document.querySelector('#onboarding'); const app = document.querySelector('#app'); onboarding.style.setProperty('display', 'none', 'important'); app.inert = false; app.removeAttribute('aria-hidden'); document.querySelector('#prompt')?.focus();
Also, warning:
# Share chats with Bullet — helps us improve model routing and answer quality
Enabled by default.
Are you freaking kidding me with YC throwing money at something like this? I guess I can fund raise just by having built https://maki.sh, and months ahead of other founders too...