← ClaudeAtlas

tilegym-adding-cutile-kernellisted

Add a new cuTile GPU kernel operator to TileGym. Covers dispatch registration in ops.py, cuTile backend implementation, __init__.py exports, test creation, and benchmark in tests/benchmark. Use when adding, creating, or implementing a new cuTile operator/kernel in TileGym, or when asking how to register a new cuTile op.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# Adding a cuTile Kernel to TileGym End-to-end workflow for adding a new operator (e.g., `my_op`) with cuTile backend. ## Execution Rules **MUST follow these rules strictly:** 1. Use TodoWrite to create the checklist below BEFORE writing any code 2. Execute steps **in order** — do NOT skip ahead or combine steps 3. Mark each todo as `completed` after finishing, `in_progress` when starting 4. If a step is not applicable (e.g., no cuTile impl), mark it `completed` with a note, do NOT silently skip 5. Each step MUST result in a file write or explicit skip decision — no silent omissions ## Instructions MUST copy this checklist to TodoWrite at the start: ``` - [ ] Step 1: Register dispatch interface in ops.py - [ ] Step 2: Implement cuTile backend - [ ] Step 3: Register in __init__.py (cutile) - [ ] Step 4: Add tests - [ ] Step 5: Add benchmark to tests/benchmark - [ ] Step 6: Verify (run pytest + lint) ``` ## Step 1: Register dispatch interface **File**: `src/tilegym/ops/ops.py` Add a `@dispatch` function — this is the **single entry point** for all backends. ```python @dispatch( "my_op", ) def my_op( input: torch.Tensor, out: Optional[torch.Tensor] = None, **kwargs: Any, ): """ Description of my_op. Args: input: Input tensor out: Optional preallocated output tensor **kwargs: Additional arguments for backend-specific configurations Returns: torch.Tensor """ raise NotImplementedError(f"my_op is not