tilegym-adding-cutile-kernellisted
Install: claude install-skill yangwhale/CloseCrab
# Adding a cuTile Kernel to TileGym
End-to-end workflow for adding a new operator (e.g., `my_op`) with cuTile backend.
## Execution Rules
**MUST follow these rules strictly:**
1. Use TodoWrite to create the checklist below BEFORE writing any code
2. Execute steps **in order** — do NOT skip ahead or combine steps
3. Mark each todo as `completed` after finishing, `in_progress` when starting
4. If a step is not applicable (e.g., no cuTile impl), mark it `completed` with a note, do NOT silently skip
5. Each step MUST result in a file write or explicit skip decision — no silent omissions
## Instructions
MUST copy this checklist to TodoWrite at the start:
```
- [ ] Step 1: Register dispatch interface in ops.py
- [ ] Step 2: Implement cuTile backend
- [ ] Step 3: Register in __init__.py (cutile)
- [ ] Step 4: Add tests
- [ ] Step 5: Add benchmark to tests/benchmark
- [ ] Step 6: Verify (run pytest + lint)
```
## Step 1: Register dispatch interface
**File**: `src/tilegym/ops/ops.py`
Add a `@dispatch` function — this is the **single entry point** for all backends.
```python
@dispatch(
"my_op",
)
def my_op(
input: torch.Tensor,
out: Optional[torch.Tensor] = None,
**kwargs: Any,
):
"""
Description of my_op.
Args:
input: Input tensor
out: Optional preallocated output tensor
**kwargs: Additional arguments for backend-specific configurations
Returns:
torch.Tensor
"""
raise NotImplementedError(f"my_op is not