I&#x27;m confused.You claim power users opt in to telemetry, and then immediately say power users opt out.

The problem with opt-in telemetry is that 95% of users are sick and tired of being spied on with every little thing they do.

If Charmin put sensors in toilet paper rolls to optimize the wiping experience, it would be dystopian. Why do we give software a pass? Privacy is a right not a telemetry problem and opt-out by default is non-consensual surveillance.

The problem with opt-in telemetry is that 95% of users don&#x27;t change defaults, and the 5% who do are your power users. They&#x27;re not representative of the average user. And only a subset of them will turn it onIronically enough the opposite happens with opt-out telemetry, for the same reason: a lot of power users will turn off telemetry, thus you will never see their usage patterns and will have to infer them. Dogfooding helps.

Fair criticism. We took a similar approach to established dev tools like Homebrew, with an anonymous, opt-out telemetry to understand install issues, crashes, and high-level usage. For cua-driver specifically, telemetry is limited to command&#x2F;tool-level events and basic environment metadata. We don’t send screenshots, recordings, app contents, prompts, typed text, file paths, or tool arguments. That said, we should make the opt-out path clearer

We don&#x27;t have a specific testing framework yet. cua-driver is closer to an automation interface than a test runner. that said, you could definitely build one on top of it. For reference these are some of our integration tests:
<a href="https:&#x2F;&#x2F;github.com&#x2F;trycua&#x2F;cua&#x2F;tree&#x2F;main&#x2F;libs&#x2F;cua-driver&#x2F;Tests&#x2F;integration" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;trycua&#x2F;cua&#x2F;tree&#x2F;main&#x2F;libs&#x2F;cua-driver&#x2F;Test...</a>One useful trick is to cua-driver &#x27;launch_app&#x27; instead of the default &#x27;open&#x27; or other osascript, since it can start the app without raising&#x2F;focusing it, and the tests don&#x27;t disturb your active desktop while they run

Would you be open to sharing what you built for running the automation tests? I could really use this right now.

Ex-Apple engineer here. I really like your implementation. A few years ago I built a similar tool to help me automate the testing of some of my native macOS apps. Being able to run multiple UI automation tests simultaneously was the big win in my case.My only criticism is enabling telemetry by default. I&#x27;m a fan of having people opt-in.

Thanks for starting that thread, I definitely drew some inspiration from it. But ultimately the secret sauce for the background click came from discovering yabai&#x27;s window_manager_focus_window_without_raise <a href="https:&#x2F;&#x2F;github.com&#x2F;asmvik&#x2F;yabai&#x2F;blob&#x2F;f17ef88116b0d988b834bb2801c3caad952cc49d&#x2F;src&#x2F;window_manager.h#L149" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;asmvik&#x2F;yabai&#x2F;blob&#x2F;f17ef88116b0d988b834bb2...</a>

Nice! Thanks for the technical writeup, ~2 weeks from me wondering how it&#x27;s implemented [1] to being able to play with a replicated version![1] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47799128">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47799128</a>

really appreciate it. macOS has powerful primitives already, but they weren’t designed as one coherent agent API so you end up stitching together and hitting roadblocks. If Apple doesn&#x27;t make this more first-class, Linux&#x2F;Android-style environments may move faster because they’re easier to instrument. I think the OpenAI&#x2F;Jony Ive AI hardware rumors are yet another signal that people may start building agent-native CUA devices instead of retrofitting agents onto existing desktops

This is one of the coolest hacks I&#x27;ve seen recently. Having done some much less involved MacOS hacking, I can&#x27;t help but wonder if we may finally see momentum behind some flavor of agent-friendly Linux&#x2F;Android if Apple doesn&#x27;t give us more ways to let agents interact with our machines.

Nothing prevents using it as a general automation library.If you want to use it directly as an automation framework, you can take a Swift dependency on &#x27;CuaDriverCore&#x27;:
<a href="https:&#x2F;&#x2F;cua.ai&#x2F;docs&#x2F;cua-driver&#x2F;guide&#x2F;getting-started&#x2F;swift-integration#cuadrivercore">https:&#x2F;&#x2F;cua.ai&#x2F;docs&#x2F;cua-driver&#x2F;guide&#x2F;getting-started&#x2F;swift-i...</a>

What is specific about this for using with agents? As opposed to offering it as a general automation library for any use?

I did something similar on Windows by creating a &quot;virtual desktop,&quot; where I can give the app focus without stealing it from another one. The idea was to basically reimplement RemoteApp without needing a dedicated Windows server.
However, in that case, the app is not visible to the user unless you use &quot;connect&quot; to the virtual desktop; to do it, I implemented (WIP) a simple VNC server in C#.

Thanks! We haven&#x27;t gone deep on Windows yet because we&#x27;re still focused on polishing the macOS release. We want to go deeper on the Mac experience before going broader across platforms, and there are still a lot of features we want to ship and use cases we want to share.

Incredible! I’m interested in doing something similar on windows, have you looked into that at all? Apparently codex computer use plans to support this on windows in the future. Were you able to see how codex was doing it, or the inspiration was just “they’ve shown it’s possible”?

Thanks for trying out Lume! We definitely haven&#x27;t given up on the idea of sandboxing GUI agents in local macOS VMs. Cua Driver is aimed at a different use case though, letting coding agents and general agents use the Mac you&#x27;re already on, asynchronously and in the background. That also makes the economics better since multiple agents can share the same machine instead of each needing its own VM

<a href="http:&#x2F;&#x2F;tart.run" rel="nofollow">http:&#x2F;&#x2F;tart.run</a> makes the VM part easy. So what if it&#x27;s overkill?

Same here. I give agents supervised direct access on my Mac for a side project. Session stealing is annoying. VM feels overkill for solo dev, but hate that the cursor jumps around while I try to do other things. Background driver sounds like the missing middle ground.

I tried out their Loom vm software a couple of months back. Worked well, fwiw. I&#x27;m not using it anymore because I decided to just give agents direct (supervised) access to my devices.

A few examples i&#x27;m excited about:- Closing the coding feedback loop by having agents verify their own changes in a real app- Automating repetitive workflows across apps that don&#x27;t have good APIs- Agents recording product demos of them using software. One compelling use case here: <a href="https:&#x2F;&#x2F;x.com&#x2F;trycua&#x2F;status&#x2F;2047383207612645426" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;trycua&#x2F;status&#x2F;2047383207612645426</a>- Creating CLI and APIs for apps by reverse implementing their GUI, e.g. see: <a href="https:&#x2F;&#x2F;github.com&#x2F;HKUDS&#x2F;CLI-Anything" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;HKUDS&#x2F;CLI-Anything</a>

Being new to the idea of using agents to run programs on one’s computer, could someone provide several use cases?

Its looking great.The audit trail question is interesting and I haven&#x27;t seen it come up much. When an agent clicks through an ERP or edits a file, you&#x27;ve got logs, but how do you explain the &quot;why&quot; behind each decision to, say, a compliance team?Curious if that&#x27;s something you&#x27;re thinking about or if it&#x27;s too early.

Show HN: Drive any macOS app in the background without stealing the cursor