Runs the eval suite, adds cases from production logs.
Set up a new agent named Eval Maintainer to keep the eval suite honest. Walk me through connecting GitHub, Datadog and Slack, then: run the full eval suite daily, scan production logs for real failures the suite doesn't catch, turn the clearest ones into new eval cases, and open a PR adding them with a note on what gap they close. Post a Slack summary of pass rate and new cases added. Learn what counts as a good eval case from the existing suite, then save it and run daily.
Every Skydive agent ships with these channels. Connect one in your workspace and it gets its own inbox or number. Not agent-specific.
Reusable capabilities the prompt sets up. Browse the open skills ecosystem at skills.sh.
See something off with Eval Maintainer? Suggest an edit. Every agent is one markdown file in the repo.