@menhguin - casual plug that @hud_evals evaluates whether AI agents can
Minh Nhat Nguyen✓@menhguin
> at local AI meetup
> dudes working on "AI agents"
> ask if they mean chat-with-internal-docs, or agents autonomously navigating, making decisions and performing complex tasks
> it's a good AI agent sir
> look inside
> chat with internal docs
image not captured
Minh Nhat Nguyen✓@menhguin
i consider it an agent if it meaningfully decides what tasks to complete, and then completes them iteratively without you fully specifying these tasks in advance
Minh Nhat Nguyen✓@menhguin
2025-05-29casual plug that @hud_evals evaluates whether AI agents can actually do various tasks
[link to Tweet](https://x.com/jayendra_ram/status/1927477447265558610)
Jay@jayendra_ram2025-05-27Over the last few months, the team at @hud_evals has made a lot of evals and environments. When we first started, we ran into a lot of problems: 1) Hosting CUA evals is annoying 2) Creating RL environments and problems is hard 3) Reviewing trajectories was super tedious 4) There was no good to write and QA evals To solve these problems, we built the hud SDK. 1/n