← All creators

jfrog

Organization

Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemini CLI, Goose, OpenCode, or any custom agent you plug in; verify behavior with rule checks, workspace diffs, multi-judge LLM consensus; pin reliability with pass^k variance across trials. Git worktrees, optional Docker sandbox.

4 indexed · 0 Featured · 18 stars · avg score 75

Categories

Indexed Skills (4)

Bio shown is the top-scored skill's repo description as a fallback — real GitHub bios land in a future update.