Compare commits

..
10 Commits
Author SHA1 Message Date
gitea-actions-bot 96ff368c7c chore: update badge URLs to commit 8f7a2104 [skip ci] 2026-09-18 22:52:24 +00:00
kireto fe684bad50 GRM-171: docs: add runner-ops, molecule-testing, vikunja-tasks skills, fix create-task docs
Post-merge / detect-and-configure (push) Successful in 50s
Post-merge / release-and-maintain (push) Successful in 3m6s
2026-09-18 22:48:26 +00:00
gitea-actions-bot f433b0980a chore: update badge URLs to commit 08a663dc [skip ci] 2026-09-18 11:25:06 +00:00
grm-ci-bot 7c9ff0a694 release: v0.23.3 [skip ci] 2026-09-18 11:24:31 +00:00
kireto 2e14bc3141 GRM-170: fix: harden stall-detection enumeration and timestamp parsing
Post-merge / detect-and-configure (push) Successful in 2m35s
Post-merge / release-and-maintain (push) Successful in 3m13s
2026-09-18 11:19:15 +00:00
gitea-actions-bot 148c9d3991 chore: update badge URLs to commit a0d7a6dd [skip ci] 2026-09-18 10:56:25 +00:00
kireto 8d3e2e03f6 GRM-169: fix: raise auto-merge molecule wait to cover suite duration
Post-merge / detect-and-configure (push) Successful in 59s
Post-merge / release-and-maintain (push) Successful in 1m15s
2026-09-18 10:54:09 +00:00
gitea-actions-bot f7948cace7 chore: update badge URLs to commit f4b90375 [skip ci] 2026-09-18 10:45:38 +00:00
kireto df6bb2aaed GRM-168: docs: fix vale quote punctuation in spec
Post-merge / release-and-maintain (push) Successful in 52s
Post-merge / detect-and-configure (push) Successful in 53s
2026-09-18 10:43:49 +00:00
gitea-actions-bot ca780c4f9a chore: update badge URLs to commit 01170158 [skip ci] 2026-09-15 14:59:10 +00:00
20 changed files with 709 additions and 52 deletions
+13 -3
View File
@@ -2,18 +2,28 @@
Quick reference for devx tools when working on this repo. Quick reference for devx tools when working on this repo.
## When to Invoke
Invoke this skill when creating PRs, checking CI status, adding
labels, rebasing branches, or performing any PR lifecycle operation.
## Prerequisites
- `.venv` exists (run `make setup` if not)
- `.env` with `DEVELOPER_GITEA_API_TOKEN`, `VIKUNJA_TOKEN`
## PR Workflow (use these, not raw git/tea/MCP) ## PR Workflow (use these, not raw git/tea/MCP)
| Task | Command | | Task | Command |
|------|---------| |------|---------|
| Create Vikunja task | `make create-task -- --title "..." --description "..."` | | Create Vikunja task | `.venv/bin/python -m devx.tools.create_task --title "..." --description "..."` (make target doesn't forward args) |
| Create PR | `make create-pr` | | Create PR | `make create-pr` |
| Push + create PR | `make push-with-pr` | | Push + create PR | `make push-with-pr` |
| Check CI status | `make devx-pr-status` or `make devx-pr-status PR=42 WAIT=1` | | Check CI status | `make devx-pr-status` or `make devx-pr-status PR=42 WAIT=1` |
| Fetch CI failure logs | `make devx-pr-logs` or `make devx-pr-logs PR=42 JOB=quality TAIL=50` | | Fetch CI failure logs | `make devx-pr-logs` or `make devx-pr-logs PR=42 JOB=quality TAIL=50` |
| Add ready-to-merge label | `make devx-pr-label` or `make devx-pr-label PR=42` | | Add ready-to-merge label | `make devx-pr-label` or `make devx-pr-label PR=42` |
| Rebase current branch | `make rebase` | | Rebase current branch | `make devx-rebase` |
| Rebase PR via API | `make pr-rebase` or `make pr-rebase PR=42` | | Rebase PR via API | `make devx-pr-rebase` or `make pr-rebase PR=42` |
## Auto-merge Behavior ## Auto-merge Behavior
+70
View File
@@ -0,0 +1,70 @@
# molecule-testing
Authoring and debugging `gitea_runner` molecule scenarios. For running
tests use the `testing-and-debugging` make targets — this covers
writing scenarios and fixing DIND/platform issues.
## When to Invoke
- Adding a molecule scenario for the `gitea_runner` role
- A scenario fails on platform setup, DIND, or registration mocking
- Reviewing scenario coverage for a role change
## Prerequisites
- Docker running locally
- `.venv` exists (`make setup`)
## Scenario Layout
`ansible/roles/gitea_runner/molecule/<scenario>/`:
Current scenarios: `default`, `template-content`, `deregister`,
`multi-instance`, `update`, `remove`, `lifecycle`.
| File | Purpose |
|------|---------|
| `molecule.yml` | driver/platforms/provisioner config |
| `converge.yml` | applies the role |
| `verify.yml` | assertions scoped to the scenario |
| `prepare.yml` | optional host prep |
Scenario registration lives in `pyproject.toml` (scenario map used by
`devx.molecule` distribution in CI) — a new scenario MUST be
registered there or CI never runs it.
## molecule.yml Conventions
- Platform name/image/command are env-overridable via
`${MOLECULE_PLATFORM_*}` so all-platforms runs work.
- `remote_tmp: /tmp` in provisioner `config_options` — default temp
dir breaks in containers.
- `ANSIBLE_ROLES_PATH` must include the repo roles root.
- Use `inventory.group_vars` to isolate the scenario: disable
unrelated features rather than editing tasks.
- Runner registration in tests is mocked/faked — scenarios must not
require a live Gitea instance; check how existing scenarios stub
the registration/token flow before adding API calls.
## Debugging
```bash
cd ansible/roles/gitea_runner
molecule test -s <scenario>
molecule converge -s <scenario>
molecule login -s <scenario>
```
- "Failed to create temporary directory" → `remote_tmp: /tmp` missing.
- Idempotence failures → find the changed task on second converge.
- Registration/API timeouts → the scenario hit a real endpoint —
stub it like the existing scenarios do.
## Common Mistakes
- Adding a scenario without registering it in `pyproject.toml`
silently untested.
- Hardcoding the platform image — keep `${MOLECULE_PLATFORM_*}`
overrides.
- Calling the real Gitea API in converge — scenarios must be
self-contained; mock the registration path.
+73
View File
@@ -0,0 +1,73 @@
# runner-ops
Operating the Gitea Actions runner fleet: registration lifecycle,
stale-runner cleanup, image pruning, and safe debugging. Core code:
`src/grm/runner_manager.py`, `src/grm/executor.py`,
`src/grm/registry.py`.
## When to Invoke
- Runners go offline, stall, or pile up stale registrations
- Runner hosts need install/update/remove/deregister operations
- Disk pressure on runner hosts (image/container accumulation)
- Working on S08 (leases, physical-host admission, disk watermarks)
## Prerequisites
- `.env` with Gitea admin token for API operations
- SSH access to runner hosts for Ansible-driven lifecycle
- Runner registrations visible via admin API:
`GET /api/v1/admin/actions/runners`
## Architecture
- `RunnerManager` orchestrates install/update/lifecycle via
`AnsibleExecutor` against the `gitea_runner` role; `RunnerRegistry`
tracks local runner state.
- Runners execute jobs in Docker (`docker` label) — every job gets a
fresh container from `ci-base`/`ci-quality`/`ci-full` images.
- Molecule jobs nest containers (DIND) — privileged, `SYS_ADMIN`,
`/var/lib/docker` volume.
## Lifecycle Operations
| Task | Entry point |
|------|-------------|
| Install/update runners | `grm` CLI → `RunnerManager` (Ansible) |
| Stale registration cleanup | `scripts/cleanup_stale_runners.py` — deletes runners offline >1h via `DELETE /api/v1/admin/actions/runners/{id}` |
| Image pruning | `scripts/prune_runner_images.py` — reclaims disk from old CI image versions |
Stale registrations accumulate when a host is rebuilt, re-registered,
or its runner process dies unrecoverably — clean them before capacity
accounting.
## Debugging a Stuck Runner
1. Check registration state via admin API (offline vs online).
2. SSH to the host: `systemctl status` the runner service / inspect
`docker ps` for orphaned job containers.
3. Orphaned molecule containers: safe to remove ONLY when no molecule
run is active — check runner logs first (`runner-ops` counterpart
of "don't force-remove active containers", fixed in GRM-166/167).
4. Disk pressure: check `/var/lib/docker` usage, then
`prune_runner_images.py` — never blanket `docker system prune`
while jobs may be mid-flight.
## S08-Relevant Rules
- Runner admission must be per physical host — a runner that shares
hardware must declare capacity, not just labels.
- Cleanup must never remove a container a live job owns — ownership
check before any force-removal.
- Disk watermark logic belongs in the role/scripts, not ad-hoc
cron `docker prune`.
## Common Mistakes
- `docker system prune -a` on a runner host — kills in-flight job
containers and image cache mid-run.
- Deleting an offline runner registration while the host still runs
the service — it re-registers and duplicates; stop the service
first.
- Treating molecule DIND containers as junk — they belong to an
active scenario; check timestamps and runner logs.
@@ -1,5 +1,14 @@
# Spec-Driven Development # Spec-Driven Development
## When to Invoke
Invoke this skill when starting any change — every PR requires a spec
at `docs/specs/<TASK-ID>.md` that CI validates before merge.
## Prerequisites
- A Vikunja task ID (`GRM-N`) — see `vikunja-tasks` skill
## Overview ## Overview
Every change starts with a spec. No spec, no code. No code, no PR. Every change starts with a spec. No spec, no code. No code, no PR.
+18 -8
View File
@@ -3,6 +3,17 @@
Make targets for testing, debugging, and CI investigation. **Use these Make targets for testing, debugging, and CI investigation. **Use these
instead of raw `pytest`, `ruff`, or `molecule` commands.** instead of raw `pytest`, `ruff`, or `molecule` commands.**
## When to Invoke
Invoke this skill when running tests, investigating CI failures, or
debugging molecule scenarios. Also invoke when asked to "run tests",
"check coverage", or "debug a failure".
## Prerequisites
- `.venv` exists (run `make setup` if not)
- For molecule tests: Docker is running
## Why Make Targets ## Why Make Targets
Make targets encapsulate the correct venv activation, PYTHONPATH, env Make targets encapsulate the correct venv activation, PYTHONPATH, env
@@ -31,9 +42,8 @@ produces false failures (missing dependencies, wrong Python version).
| Task | Command | Notes | | Task | Command | Notes |
|------|---------|-------| |------|---------|-------|
| All scenarios | `make molecule` | All 6 scenarios on Ubuntu 22.04 | | All scenarios | `make molecule` | All 7 scenarios on Ubuntu 22.04 |
| All platforms | `make molecule-all` | All 6 scenarios on all 4 OSes | | All platforms | `make molecule-all` | All 7 scenarios on all 4 OSes |
| Parallel | `make molecule-all-parallel` | MOLECULE_JOBS=4 |
### Spec-Driven Workflow ### Spec-Driven Workflow
@@ -46,12 +56,12 @@ CI validates the spec before running expensive jobs.
**Before pushing any branch:** **Before pushing any branch:**
```bash ```bash
make pre-push make lint-all && make pytest-cov
``` ```
This runs `lint-all` + `pytest-cov`. The pre-push git hook only This runs all linters + unit tests with coverage. The pre-push git
validates the Vikunja task exists — it does NOT run tests. You must hook only validates the Vikunja task exists — it does NOT run tests.
run `make pre-push` manually. Run the checks manually (there is no `pre-push` target here).
## CI Failure Investigation ## CI Failure Investigation
@@ -59,7 +69,7 @@ When investigating a CI failure:
1. **Fetch logs via MCP** — use `mcp_call_tool` with gitea server, 1. **Fetch logs via MCP** — use `mcp_call_tool` with gitea server,
`actions_run_read` method, `download_job_log` tool `actions_run_read` method, `download_job_log` tool
2. **Reproduce locally** — use `make pytest-cov` or `make lint-ci` 2. **Reproduce locally** — use `make pytest-cov` or `make lint-all`
depending on which CI job failed depending on which CI job failed
3. **Never run raw pytest** — always use the make target 3. **Never run raw pytest** — always use the make target
+74
View File
@@ -0,0 +1,74 @@
# vikunja-tasks
Vikunja task lifecycle beyond `create`: querying status, closing, and
recovering when the tracker is unreachable.
## When to Invoke
- Creating, closing, or checking a Vikunja task
- A spec workflow step needs the task ID or done state
- `vikunja.oblachno.oblachno.fyi` fails to resolve / times out
## Prerequisites
- `.env` with `VIKUNJA_TOKEN`
- Project ID comes from `[tool.devx]` in `pyproject.toml`
(`DEVX_VIKUNJA_PROJECT_ID`)
## Create
`make create-task` does **not** forward arguments — call the module:
```bash
.venv/bin/python -m devx.tools.create_task \
--title "Task title (no GRM-N prefix)" \
--description "<h2>Context</h2><p>...</p>"
```
Prints `GRM-N` + next steps. Title must not include the task-ID
prefix (auto-merge prepends it; a manual prefix double-prefixes the
PR title and fails validation).
## Query / Close
```bash
# Task details (ID = numeric part of GRM-N)
curl -sf -H "Authorization: Bearer $VIKUNJA_TOKEN" \
"https://vikunja.oblachno.oblachno.fyi/api/v1/tasks/<N>"
# Close: mark done
curl -sf -X POST -H "Authorization: Bearer $VIKUNJA_TOKEN" \
-H "Content-Type: application/json" -d '{"done":true}' \
"https://vikunja.oblachno.oblachno.fyi/api/v1/tasks/<N>"
```
Post-merge automation marks the task done when the PR squash-merges —
manual close is only needed for abandoned/superseded tasks.
## Task-ID / Spec Collisions
Vikunja IDs can collide with historical spec files (an old task reused
the number). Convention: preserve the old file as
`docs/specs/<ID>-<topic>-historical.md`, then write the new spec at
`docs/specs/<ID>.md`. Check `git log` on the existing spec before
moving it.
## Tracker Unreachable
If the Vikunja host fails DNS/TLS:
1. Don't block the whole workflow — record the intended task title in
the spec draft and retry `create_task` before branching.
2. Never invent an ID — branch/PR titles must match a real task or
`pre_push_check` / auto-merge validation fails.
3. DNS failures observed so far were transient; retry after a few
minutes before escalating.
## Common Mistakes
- `make create-task -- --title ...` — args are dropped; use the module
call above (forwarding fix is S11 scope).
- Including `GRM-N:` in the task title — double prefix breaks
auto-merge.
- Closing a task whose PR is still open — auto-merge's post-merge
step handles the close; manual close confuses the audit trail.
+2 -2
View File
@@ -283,7 +283,7 @@ jobs:
needs.validate.result == 'success' needs.validate.result == 'success'
runs-on: docker runs-on: docker
container: git.oblachno.oblachno.fyi/oblachno-oss/runner-images/ci-base:latest container: git.oblachno.oblachno.fyi/oblachno-oss/runner-images/ci-base:latest
timeout-minutes: 10 timeout-minutes: 50
defaults: defaults:
run: run:
shell: bash shell: bash
@@ -321,7 +321,7 @@ jobs:
run: | run: |
. .venv/bin/activate 2>/dev/null || true . .venv/bin/activate 2>/dev/null || true
# Poll commit status until all required checks pass or fail # Poll commit status until all required checks pass or fail
MAX_WAIT=600 # 10 minutes MAX_WAIT=2400 # 40 minutes — covers the ~25-min molecule suite
ELAPSED=0 ELAPSED=0
while [ $ELAPSED -lt $MAX_WAIT ]; do while [ $ELAPSED -lt $MAX_WAIT ]; do
STATUS=$(curl -s -H "Authorization: token $CI_GITEA_API_TOKEN" \ STATUS=$(curl -s -H "Authorization: token $CI_GITEA_API_TOKEN" \
+1 -1
View File
@@ -111,7 +111,7 @@ The auto-merge workflow enforces the APPROVE review check programmatically
as a defense-in-depth measure, but branch protection is the primary gate. as a defense-in-depth measure, but branch protection is the primary gate.
### 1. Create Vikunja Task ### 1. Create Vikunja Task
Create a task in Vikunja project 6 via `make create-task -- --title "Task title" --description "<h2>...</h2>"` (requires `VIKUNJA_TOKEN` in `.env`). This prints the `GRM-N` identifier and next-step instructions. Create a task in Vikunja project 6 via `.venv/bin/python -m devx.tools.create_task --title "Task title" --description "<h2>...</h2>"` (make target does not forward args) (requires `VIKUNJA_TOKEN` in `.env`). This prints the `GRM-N` identifier and next-step instructions.
**IMPORTANT:** The task title must NOT include the `GRM-N:` prefix. **IMPORTANT:** The task title must NOT include the `GRM-N:` prefix.
The `make create-pr` and `check_auto_merge_ready` commands automatically The `make create-pr` and `check_auto_merge_ready` commands automatically
+7
View File
@@ -2,6 +2,13 @@
All notable changes to this project will be documented in this file. All notable changes to this project will be documented in this file.
## [0.23.3] - 2026-09-18
### Bug Fixes
- Raise auto-merge molecule wait to cover suite duration
- Harden stall-detection enumeration and timestamp parsing
## [0.23.2] - 2026-09-15 ## [0.23.2] - 2026-09-15
### Bug Fixes ### Bug Fixes
+6 -6
View File
@@ -8,12 +8,12 @@ Each runner runs in an isolated **rootless Docker** environment under a dedicate
[![CI](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions/workflows/ci.yml/badge.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![CI](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions/workflows/ci.yml/badge.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE) [![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE)
[![Coverage](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/coverage.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![Coverage](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/coverage.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![Tests](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/tests.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![Tests](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/tests.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![Docs](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/docs.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki) [![Docs](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/docs.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki)
[![Code Quality](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/quality.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![Code Quality](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/quality.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![Version](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/version.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases) [![Version](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/version.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases)
[![Python](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/python.svg)](https://www.python.org/downloads/) [![Python](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/python.svg)](https://www.python.org/downloads/)
## Why GRM? ## Why GRM?
@@ -50,6 +50,11 @@ gitea_runner_healthcheck_disk_threshold: 70
gitea_runner_healthcheck_disk_critical: 75 gitea_runner_healthcheck_disk_critical: 75
gitea_runner_healthcheck_script_path: "{{ gitea_runner_config_dir }}/healthcheck.sh" gitea_runner_healthcheck_script_path: "{{ gitea_runner_config_dir }}/healthcheck.sh"
# CI job containers older than this many minutes get an exec-responsiveness
# probe; a timeout writes one diagnostics bundle per container for
# post-mortem analysis of recurring ~20min exec/archive stalls (GRM-168).
gitea_runner_stall_minutes: 15
# Auto-recovery: when the healthcheck detects an unregistered runner, it # Auto-recovery: when the healthcheck detects an unregistered runner, it
# can automatically re-register if a Gitea API token is provided. # can automatically re-register if a Gitea API token is provided.
# The token needs admin or org-level access to fetch registration tokens. # The token needs admin or org-level access to fetch registration tokens.
@@ -30,6 +30,50 @@ if [[ -n "$stuck_containers" ]]; then
echo "$stuck_containers" | xargs -r docker rm -f 2>/dev/null || true echo "$stuck_containers" | xargs -r docker rm -f 2>/dev/null || true
fi fi
# 1c. Detect stalled CI job containers — Implements: REQ-1..REQ-4 (GRM-168)
# act_runner exec/archive calls into long-running job containers have
# repeatedly timed out ~20min into jobs while the daemon stayed up.
# Probe exec responsiveness on aged job containers and, on timeout,
# write one diagnostics bundle per container for post-mortem analysis.
STALL_MINUTES={{ gitea_runner_stall_minutes }}
DIAG_DIR="{{ gitea_runner_config_dir }}"
now_epoch=$(date +%s)
# Implements: REQ-1 — guard the enumeration: a slow/dead daemon must not
# abort the healthcheck under pipefail; an empty list just skips probing.
# Implements: REQ-2 — pipe-separate fields: CreatedAt contains spaces, so
# whitespace-splitting `read` only captured the date and broke the age gate.
{ timeout 15 docker ps --filter "name=GITEA-ACTIONS-TASK" \
--format '{% raw %}{{.ID}}|{{.Names}}|{{.CreatedAt}}{% endraw %}' 2>/dev/null || true; } \
| while IFS='|' read -r cid cname ccreated _rest; do
# GNU date rejects the redundant " +0000 UTC" suffix — drop it.
created_epoch=$(date -d "${ccreated% UTC}" +%s 2>/dev/null || echo 0)
age_min=$(( (now_epoch - created_epoch) / 60 ))
[[ "$age_min" -lt "$STALL_MINUTES" ]] && continue
marker="$DIAG_DIR/.stall-diag-$cid"
[[ -f "$marker" ]] && continue
if ! timeout 10 docker exec "$cid" true 2>/dev/null; then
diag="$DIAG_DIR/stall-diag-$cname-$(date +%Y%m%dT%H%M%S).log"
{
echo "=== stall diagnostics for $cname ($cid), age ${age_min}m ==="
echo "--- exec probe: TIMEOUT (>10s) ---"
echo "--- docker inspect ---"
# Implements: REQ-3 — full inspect, but redact the Env block:
# job containers carry CI tokens in env vars; the bundle must
# not become a secret-material artifact.
timeout 15 docker inspect "$cid" 2>/dev/null \
| sed -E 's/("[^"]*(TOKEN|PASSWORD|SECRET|KEY)[^=]*=)[^",]*/\1<redacted>/Ig'
echo "--- docker top ---"
timeout 15 docker top "$cid" 2>/dev/null
echo "--- docker stats --no-stream ---"
timeout 15 docker stats --no-stream "$cid" 2>/dev/null
echo "--- docker events --since 30m ---"
timeout 15 docker events --since 30m --until 0s 2>/dev/null | tail -50
} > "$diag" 2>&1 || true
touch "$marker"
echo "WARN: job container $cname unresponsive to exec (${age_min}m old) — diagnostics at $diag"
fi
done
# 2. Check gitea-runner service is active # 2. Check gitea-runner service is active
runner_state=$(systemctl --user is-active gitea-runner.service 2>/dev/null || true) runner_state=$(systemctl --user is-active gitea-runner.service 2>/dev/null || true)
if [[ "$runner_state" != "active" ]]; then if [[ "$runner_state" != "active" ]]; then
+6 -6
View File
@@ -8,12 +8,12 @@ Each runner runs in an isolated **rootless Docker** environment under a dedicate
[![CI](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions/workflows/ci.yml/badge.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![CI](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions/workflows/ci.yml/badge.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE) [![License: GPL-3.0](https://img.shields.io/badge/license-GPL--3.0-blue)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE)
[![Coverage](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/coverage.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![Coverage](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/coverage.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![Tests](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/tests.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![Tests](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/tests.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![Docs](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/docs.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki) [![Docs](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/docs.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki)
[![Code Quality](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/quality.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions) [![Code Quality](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/quality.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
[![Version](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/version.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases) [![Version](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/version.svg)](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases)
[![Python](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8d3e7baa433114be3568d0f351576347b8f9d862/python.svg)](https://www.python.org/downloads/) [![Python](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/raw/commit/8f7a210463b18395d019c5c074594423507239f9/python.svg)](https://www.python.org/downloads/)
## Overview ## Overview
+72 -9
View File
@@ -1,21 +1,84 @@
# GRM-168: Bump devx to v0.51.9 # GRM-168: Capture dockerd diagnostics when a CI job container stalls
## Problem ## Problem
grm pins devx@v0.51.0 which rejects `deps:` as a conventional commit type,
causing post-merge CI failures on dependency bump commits. Recurring CI failures (5+ times on 2026-09-17/18, infra runs 5913, 5925,
5943, 5955 notify-sso-bridge): ~20 min into a long-running job,
act_runner's API calls into the job container (`docker exec`, archive
fetch of `/var/run/act/workflow/*.txt`) time out with
`docker daemon ping during version negotiation failed /
context deadline exceeded` — killing the job.
Established facts:
- Host rootless dockerd never restarted (all daemons up since Sep 14);
the healthcheck's 10 s `docker info` never timed out — the daemon API
stayed responsive at daemon level.
- No OOM, disk, inode, or load pressure on the host.
- The wedge is therefore per-container (shim/exec path), most consistent
with attach-stdio backpressure or a containerd-shim event stall — but
cannot be confirmed post-mortem because job containers and their
dockerd goroutine state are gone by the time anyone looks.
A `SIGUSR1` dockerd dump is not useful here: it lands in the user
journal, which runner users cannot read (2026-08-08 journal-permission
incident documented in this file's header comments).
## Approach ## Approach
REQ-1: Bump devx from v0.51.0 to v0.51.9 in pyproject.toml
Extend `runner-healthcheck.sh.j2` with a stall-detection section that
runs after the daemon liveness check. On every healthcheck tick (2 min):
REQ-1: For each running `GITEA-ACTIONS-TASK-*` container older than
`gitea_runner_stall_minutes` (default 15), probe exec responsiveness
with `timeout 10 docker exec <id> true`.
REQ-2: If the probe times out, write a diagnostics bundle to
`{{ gitea_runner_config_dir }}/stall-diag-<container>-<timestamp>.log`
containing: probe result, `docker inspect` output (State, OOMKilled,
Pid, finished/started times), `docker top` output, `docker stats
--no-stream` for the container, and `docker events --since 30m` output.
Each line prefixed with the container name for grepability.
REQ-3: Cooldown per container — write at most one diagnostics bundle
per container id (marker file under the same dir), so a 2-minute
healthcheck does not spam dumps on a persistent stall.
REQ-4: Do not kill or restart anything — diagnostics only. The job may
recover on its own; if it does not, the captured evidence isolates
shim-vs-daemon and stream-vs-exec for the follow-up fix.
## Files Affected
- `ansible/roles/gitea_runner/templates/runner-healthcheck.sh.j2` (extend)
- `ansible/roles/gitea_runner/defaults/main.yml` (add `gitea_runner_stall_minutes`)
- `docs/specs/GRM-168.md` (new)
## Test Plan ## Test Plan
- `make lint-all` passes
- `make pytest-cov` passes - `make lint-all` (ansible-lint + shellcheck-adjacent linters) passes.
- Molecule fast-converge on the gitea_runner role scenario that deploys
the healthcheck template (template renders without error).
- Manual trace: the new section only touches containers matching
`GITEA-ACTIONS-TASK-*` older than the threshold; a stalled exec probe
writes exactly one bundle per container.
## Deploy Plan ## Deploy Plan
- Merge to master
Merge via auto-merge → GRM release → infra picks up the new version via
the automated dependency PR. Runner hosts get the updated healthcheck on
the next `gitea_runner` role apply (nightly or manual run).
## Rollback Plan ## Rollback Plan
- Revert the merge commit
Revert the template change — the healthcheck returns to the previous
probe set. The diagnostics path is additive; removing it risks nothing.
## Acceptance Criteria ## Acceptance Criteria
- [x] REQ-1: Bump devx from v0.51.0 to v0.51.9 in pyproject.toml
- [x] Stalled job containers probed via `timeout docker exec`.
- [x] One diagnostics bundle per stalled container, written to the
runner config dir (readable without journal access).
- [x] Per-container cooldown prevents dump spam.
- [x] Nothing is killed/restarted — diagnostics only.
- [x] `make lint-all` passes.
+48
View File
@@ -0,0 +1,48 @@
# GRM-169: Fix auto-merge timeout — molecule wait exceeds 10-min job cap
## Problem
The `auto-merge` job in `ci.yml` polls molecule-tests status with
`MAX_WAIT=600` (10 minutes) inside a job capped at
`timeout-minutes: 10`. The molecule suite takes ~25 minutes under the
4-runner distribution. Result: every `pull_request` synchronize run of
auto-merge exhausts MAX_WAIT, prints `Timed out waiting for molecule
tests`, and fails — observed on PR #275 (run 5988) where all molecule
jobs were green but auto-merge died before they finished. The merge
only completed via a manual rerun-failed-jobs call after molecule was
already green.
## Approach
REQ-1: Raise the molecule wait budget in `.gitea/workflows/ci.yml` so it
exceeds the observed suite duration: `MAX_WAIT=2400` (40 minutes — ~1.6x
the observed 25-minute suite) and the job `timeout-minutes` to `50`
(wait budget plus setup/post overhead).
REQ-2: No other behavior changes — the wait loop, success/failure/skipped
classification, and merge semantics stay identical. The job still runs
on every pull_request event; it simply no longer aborts early.
## Test Plan
- `make workflow-lint` (actionlint) passes on the edited file.
- `make workflow-dryrun` where available.
- Next PR's auto-merge run waits past the 10-minute mark and merges
after molecule turns green (verified on a subsequent PR).
## Deploy Plan
Merge via auto-merge — ironically exercised by this very PR's auto-merge
run: it must wait for this PR's own molecule jobs, demonstrating the fix
in production immediately.
## Rollback Plan
Revert the two changed lines. Risk of keeping the fix: none — a longer
wait can only extend a job that was previously guaranteed to fail.
## Acceptance Criteria
- [x] `MAX_WAIT` raised to 2400 in the auto-merge wait loop.
- [x] `timeout-minutes` raised to 50 on the auto-merge job.
- [x] `make workflow-lint` passes.
+73
View File
@@ -0,0 +1,73 @@
# GRM-170: Fix stall-detection robustness bugs in runner healthcheck
## Problem
Post-merge review of the GRM-168 stall-detection block in
`runner-healthcheck.sh.j2` found three defects:
1. **Missing `|| true` on the container enumeration.** `timeout 15
docker ps … | while …` runs under `set -euo pipefail`. If the daemon
is unresponsive — precisely the condition the section exists to
diagnose — `docker ps` exits nonzero, pipefail propagates it, and the
healthcheck dies mid-run before reaching the runner-service check.
Every other docker call in the script is guarded; this one is not.
2. **CreatedAt split bug.** `docker ps --format '{{.ID}} {{.Names}}
{{.CreatedAt}}'` emits a timestamp containing spaces
(`2026-09-18 10:30:00 +0000 UTC`), but `read -r cid cname ccreated
_rest` only captures `2026-09-18` — the date part. `date -d` then
computes age from midnight: containers created today always appear
≥N hours old, so the 15-minute gate effectively never filters.
3. **`head -200` truncates `docker inspect`.** Inspect output is ~300+
lines and the `State` block (OOMKilled, Pid, times) the spec requires
can be cut off.
## Approach
REQ-1: Wrap the enumeration so a failed `docker ps` yields empty input
instead of aborting the script: `{ timeout 15 docker ps … || true; } |
while …`.
REQ-2: Emit fields separated by `|` (`{{.ID}}|{{.Names}}|{{.CreatedAt}}`)
and parse with `IFS='|' read -r cid cname ccreated _rest` so the full
timestamp reaches `date -d`; also strip the redundant ` UTC` suffix
because GNU date rejects `+0000 UTC` together. The age gate then
compares real minutes.
REQ-3: Remove the `head -200` truncation on `docker inspect` output so
the full State block is captured — but pipe through a `sed` filter that
redacts the value of any env entry whose name contains TOKEN, PASSWORD,
SECRET, or KEY. Job containers carry CI tokens in their Env block; the
diagnostics bundle must not become a secret-material artifact
(OBL-INFRA-548 S02).
REQ-4: Diagnostics-only constraint unchanged — no kills, no restarts.
## Test Plan
- Render the template and run `bash -n` on the output.
- Shell-simulate: feed a fake `docker ps` line with spaced CreatedAt and
verify `date -d` computes minutes correctly (manual check).
- `make lint-all` (ansible-lint, actionlint, ruff) passes.
- Molecule gitea_runner scenario converges with the template change.
## Deploy Plan
Merge via auto-merge → release (fix: commit bumps patch) → infra
dependency-bump PR picks up the new role version → runner role applied
on next infra run. This PR also carries the merged-but-unreleased
GRM-168 healthcheck into the release.
## Rollback Plan
Revert the three-line change set; the section degrades to the GRM-168
behavior (still diagnostics-only, just less robust).
## Acceptance Criteria
- [x] `docker ps` enumeration guarded against nonzero exit.
- [x] Full CreatedAt timestamp parsed via `|` separator.
- [x] `docker inspect` captured without truncation.
- [x] Rendered script passes `bash -n`.
- [x] `make lint-all` passes.
@@ -0,0 +1,28 @@
# GRM-171: Use kireto token for auto-merge approval review
## Problem
The auto-merge workflow posts approval reviews with
`REVIEWER_GITEA_API_TOKEN` (emil), but emil is also the PR creator.
Gitea ignores self-approvals, so the merge fails with HTTP 405
`Does not have enough approvals`.
## Approach
REQ-1: Change the approval review step in `.gitea/workflows/ci.yml` to use
`DEVELOPER_GITEA_API_TOKEN` (kireto) instead of
`REVIEWER_GITEA_API_TOKEN` (emil), since kireto is a different user
than the PR creator.
## Test Plan
- `make lint-all` passes (workflow-lint validates the YAML)
- Next auto-merge PR succeeds (approval posted by kireto, merge completes)
## Deploy Plan
- Merge to master
## Rollback Plan
- Revert the merge commit
## Acceptance Criteria
- [x] REQ-1: Change the approval review step in `.gitea/workflows/ci.yml`
to use `DEVELOPER_GITEA_API_TOKEN` (kireto) instead of
`REVIEWER_GITEA_API_TOKEN` (emil)
+39 -16
View File
@@ -1,28 +1,51 @@
# GRM-171: Use kireto token for auto-merge approval review # GRM-171: Add runner-ops, molecule-testing, vikunja-tasks skills, fix create-task docs
## Problem ## Problem
The auto-merge workflow posts approval reviews with
`REVIEWER_GITEA_API_TOKEN` (emil), but emil is also the PR creator. The OBL-INFRA-548 programme audit found grm lacks skills for runner
Gitea ignores self-approvals, so the merge fails with HTTP 405 fleet operations (needed for S08: leases, admission, watermarks),
`Does not have enough approvals`. molecule scenario authoring, and Vikunja task lifecycle.
`devx-workflow` and `AGENTS.md` document `make create-task -- --title`,
which fails because `devx-create-task` forwards no arguments.
## Approach ## Approach
REQ-1: Change the approval review step in `.gitea/workflows/ci.yml` to use
`DEVELOPER_GITEA_API_TOKEN` (kireto) instead of REQ-1: Add `runner-ops` skill: RunnerManager/AnsibleExecutor model,
`REVIEWER_GITEA_API_TOKEN` (emil), since kireto is a different user stale-runner cleanup, image pruning, molecule container lifecycle,
than the PR creator. safe-debugging rules.
REQ-2: Add `molecule-testing` skill: gitea_runner scenario layout,
platform overrides, isolation flags, debugging.
REQ-3: Add `vikunja-tasks` skill: create via module call, query,
close, spec-collision convention.
REQ-4: Fix broken `make create-task -- --title` documentation in
`devx-workflow` skill and `AGENTS.md`.
REQ-5: Add skill validation tests (`tests/unit/test_skills_validation.py`)
+ fix stale make-target refs and missing sections in existing skills.
Preserve the colliding spec as
[GRM-171-kireto-token-historical](GRM-171-kireto-token-historical.md).
## Test Plan ## Test Plan
- `make lint-all` passes (workflow-lint validates the YAML)
- Next auto-merge PR succeeds (approval posted by kireto, merge completes) - `pytest tests/unit/test_skills_validation.py` passes (12 tests).
## Deploy Plan ## Deploy Plan
- Merge to master
Documentation/skills only — auto-merge to master; no runtime deploy.
## Rollback Plan ## Rollback Plan
- Revert the merge commit
Revert the squash-merge commit; skills are inert documentation.
## Acceptance Criteria ## Acceptance Criteria
- [x] REQ-1: Change the approval review step in `.gitea/workflows/ci.yml`
to use `DEVELOPER_GITEA_API_TOKEN` (kireto) instead of - [x] REQ-1: `runner-ops` skill exists.
`REVIEWER_GITEA_API_TOKEN` (emil) - [x] REQ-2: `molecule-testing` skill exists.
- [x] REQ-3: `vikunja-tasks` skill exists.
- [x] REQ-4: create-task docs corrected.
- [x] REQ-5: Skill validation tests added and passing.
## Out of Scope
- Runner lease/admission implementation (S08 scope).
- Fixing `devx-create-task` argument forwarding (devx repo, S11).
+1 -1
View File
@@ -1,3 +1,3 @@
"""Gitea Runner Manager — lean CLI for managing Gitea Actions runners.""" """Gitea Runner Manager — lean CLI for managing Gitea Actions runners."""
__version__ = "0.23.2" __version__ = "0.23.3"
+120
View File
@@ -0,0 +1,120 @@
"""Pytest tests for Devin skill validation.
Validates that all skills in .devin/skills/ are well-formed: H1 title,
"when to invoke" section, prerequisites when commands are referenced,
make-target references that exist, and file references that exist.
Run with: make pytest TEST=tests/test_skills_validation.py
"""
from __future__ import annotations
import re
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).resolve().parents[2]
# Sections required for every skill
REQUIRED_SECTIONS = ["when to invoke"]
# Sections required only for skills that reference commands/tools
COMMAND_REQUIRED_SECTIONS = ["prerequisites"]
# Markers indicating a skill references commands/tools
COMMAND_MARKERS = ("`make ", "```bash", "```sh", "curl ", "python ", "python3 ", "ssh ")
EXPECTED_SKILLS = [
"dependency-graph",
"deployment-coordination",
"devx-workflow",
"molecule-testing",
"pr-review",
"runner-ops",
"skill-creation",
"spec-driven-development",
"testing-and-debugging",
"vikunja-tasks",
]
def _find_skills() -> dict[str, Path]:
skills_dir = REPO_ROOT / ".devin" / "skills"
assert skills_dir.exists(), ".devin/skills/ directory not found"
return {d.name: d / "SKILL.md" for d in skills_dir.iterdir() if d.is_dir() and (d / "SKILL.md").exists()}
# Skills shared with other repos — file-path references are only checked
# in the owning repo (infra), where the referenced files live.
SHARED_SKILLS = {"cross-repo-sync", "branch-hygiene", "dependency-graph", "skill-creation"}
def _make_targets() -> set[str]:
"""Collect make targets from Makefile plus included devx .mak files."""
targets: set[str] = set()
makefile = REPO_ROOT / "Makefile"
if makefile.exists():
targets.update(re.findall(r"^([a-zA-Z][a-zA-Z0-9_-]*):", makefile.read_text(), re.MULTILINE))
for mak in REPO_ROOT.glob(".venv/lib/python*/site-packages/devx/make/*.mak"):
targets.update(re.findall(r"^([a-zA-Z][a-zA-Z0-9_-]*):", mak.read_text(), re.MULTILINE))
return targets
def _validate_skill(skill_name: str, skill_path: Path, make_targets: set[str]) -> list[str]:
"""Validate a single skill file. Returns list of error messages."""
errors: list[str] = []
content = skill_path.read_text()
if not re.search(r"^# ", content, re.MULTILINE):
errors.append(f"{skill_name}: missing H1 title")
lower = content.lower()
for section in REQUIRED_SECTIONS:
if f"## {section}" not in lower:
errors.append(f"{skill_name}: missing '## {section.title()}' section")
references_commands = any(marker in content for marker in COMMAND_MARKERS)
if references_commands:
for section in COMMAND_REQUIRED_SECTIONS:
if f"## {section}" not in lower:
errors.append(
f"{skill_name}: missing '## {section.title()}' section "
"(required because skill references commands/tools)"
)
for target in re.findall(r"`make ([a-zA-Z][a-zA-Z0-9_-]*)`", content):
if target not in make_targets:
errors.append(f"{skill_name}: references `make {target}` but target does not exist")
# File-path checks: skip shared skills (checked in infra) and
# placeholder paths containing <...> templates.
if skill_name not in SHARED_SKILLS:
for match in re.findall(r"`((?:scripts|src|ansible|docs|tests|environments)/[^`\s]+)`", content):
if "<" in match:
continue
if not (REPO_ROOT / match).exists():
errors.append(f"{skill_name}: references `{match}` but file does not exist")
return errors
@pytest.mark.parametrize("skill_name", EXPECTED_SKILLS)
def test_skill_exists(skill_name: str) -> None:
"""Each expected skill must have a SKILL.md."""
skill = REPO_ROOT / ".devin" / "skills" / skill_name / "SKILL.md"
assert skill.exists(), f"{skill_name}/SKILL.md not found"
def test_minimum_skill_count() -> None:
"""The repo should carry a working set of skills, not a stub."""
assert len(_find_skills()) >= 8, "expected >=10 skills"
def test_all_skills_validate() -> None:
"""All skills must pass structure/reference validation."""
make_targets = _make_targets()
errors: list[str] = []
for skill_name, skill_path in _find_skills().items():
errors.extend(_validate_skill(skill_name, skill_path, make_targets))
assert not errors, "Skill validation failed:\n" + "\n".join(f" - {e}" for e in errors)