Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f60b84c52d |
@@ -1,135 +0,0 @@
|
|||||||
# dependency-graph
|
|
||||||
|
|
||||||
Map of the oblachno ecosystem. Knows which repo produces what, which
|
|
||||||
repos depend on which, and the correct order for cross-repo changes.
|
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
Invoke this skill when:
|
|
||||||
- Changes span multiple repos
|
|
||||||
- A change in one repo requires version bumps in downstream repos
|
|
||||||
- Deploying infrastructure that depends on published packages/images
|
|
||||||
- Verifying the ecosystem is in a consistent state before deployment
|
|
||||||
- Determining which repos to update and in what order
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- All repos cloned under `/home/emo/dev/ideas/oblachno/`
|
|
||||||
- `.env` with `DEVELOPER_GITEA_API_TOKEN` in each repo
|
|
||||||
|
|
||||||
## Ecosystem Map
|
|
||||||
|
|
||||||
```
|
|
||||||
devx (PyPI package)
|
|
||||||
/ | \
|
|
||||||
/ | \
|
|
||||||
grm sso-bridge infra
|
|
||||||
(PyPI) (PyPI+Docker) (deploys all)
|
|
||||||
| | |
|
|
||||||
v v v
|
|
||||||
infra bump infra bump staging
|
|
||||||
(auto PR) (auto PR) production
|
|
||||||
|
|
|
||||||
mattermost-oidc (Docker image)
|
|
||||||
(infra pulls :latest at deploy)
|
|
||||||
```
|
|
||||||
|
|
||||||
## Repositories
|
|
||||||
|
|
||||||
| Repo | Produces | Consumers | Release Trigger |
|
|
||||||
|------|----------|-----------|-----------------|
|
|
||||||
| `devx` | PyPI package `devx` | grm, sso-bridge, infra | User-facing changes to `src/devx/**` |
|
|
||||||
| `grm` | PyPI package `grm` | infra | User-facing changes to `src/grm/**` or `ansible/**` |
|
|
||||||
| `sso-bridge` | PyPI package `sso_bridge` + Docker image | infra | User-facing changes to `src/sso_bridge/**` or `ansible/**` |
|
|
||||||
| `infra` | Staging/production deployment | (end users) | User-facing changes + nightly gate |
|
|
||||||
| `mattermost-oidc` | Docker image `mattermost-oidc` | infra (pulls at deploy) | `Dockerfile` or `build.yml` changes |
|
|
||||||
|
|
||||||
## Dependency Chain
|
|
||||||
|
|
||||||
### devx → all repos
|
|
||||||
|
|
||||||
devx publishes to the Gitea PyPI registry. grm, sso-bridge, and infra
|
|
||||||
pin devx in `pyproject.toml`:
|
|
||||||
```toml
|
|
||||||
"devx @ git+https://git.oblachno.oblachno.fyi/oblachno-oss/devx.git@vX.Y.Z"
|
|
||||||
```
|
|
||||||
|
|
||||||
When devx publishes a new version:
|
|
||||||
1. grm, sso-bridge, and infra must bump their pinned devx version
|
|
||||||
2. This is currently manual — no auto-dependency-PR from devx
|
|
||||||
3. Each repo must `make setup` to pick up the new version
|
|
||||||
|
|
||||||
### grm → infra
|
|
||||||
|
|
||||||
grm publishes to PyPI. Its post-merge workflow auto-creates an infra
|
|
||||||
dependency PR via `devx.ci.create_dependency_pr --repo oblachno/infra
|
|
||||||
--package grm`. The PR bumps the pinned grm version in infra's
|
|
||||||
`pyproject.toml`.
|
|
||||||
|
|
||||||
### sso-bridge → infra
|
|
||||||
|
|
||||||
sso-bridge publishes to PyPI AND builds a Docker image. Its post-merge
|
|
||||||
workflow auto-creates an infra dependency PR via
|
|
||||||
`devx.ci.create_dependency_pr --repo oblachno/infra --package
|
|
||||||
sso_bridge`. The PR bumps the pinned sso_bridge version.
|
|
||||||
|
|
||||||
The Docker image is pulled by infra at deploy time (`sso-bridge:latest`).
|
|
||||||
|
|
||||||
### mattermost-oidc → infra
|
|
||||||
|
|
||||||
mattermost-oidc builds a Docker image tagged `:latest` and `:MM_VERSION`.
|
|
||||||
infra pulls `mattermost-oidc:latest` at deploy time. There is no
|
|
||||||
auto-dependency-PR — infra simply pulls the latest image.
|
|
||||||
|
|
||||||
### infra → staging/production
|
|
||||||
|
|
||||||
infra deploys to staging and production. The deployment:
|
|
||||||
1. Provisions VMs from golden images
|
|
||||||
2. Runs Ansible roles (including grm and sso-bridge roles)
|
|
||||||
3. Pulls Docker images (sso-bridge, mattermost-oidc)
|
|
||||||
4. Configures services
|
|
||||||
|
|
||||||
## Correct Order for Cross-Repo Changes
|
|
||||||
|
|
||||||
When a change spans multiple repos, follow this order:
|
|
||||||
|
|
||||||
1. **devx first** — if the change starts in devx, merge and publish devx
|
|
||||||
first. Wait for the PyPI publish job to complete.
|
|
||||||
2. **Bump devx in consumers** — in grm/sso-bridge/infra, bump the pinned
|
|
||||||
devx version, run `make setup`, verify tests pass, merge.
|
|
||||||
3. **grm/sso-bridge second** — merge and publish grm/sso-bridge. Wait
|
|
||||||
for the PyPI publish + Docker image build to complete.
|
|
||||||
4. **Auto-dependency-PRs** — grm/sso-bridge post-merge auto-creates infra
|
|
||||||
PRs to bump pinned versions. Wait for these PRs to appear.
|
|
||||||
5. **Merge infra dependency PRs** — review and merge the auto-created
|
|
||||||
infra PRs.
|
|
||||||
6. **infra last** — deploy to staging, validate, promote to production.
|
|
||||||
|
|
||||||
## State Verification Before Deployment
|
|
||||||
|
|
||||||
Before deploying infra, verify:
|
|
||||||
|
|
||||||
1. **devx version consistent** — all repos pin the same devx version
|
|
||||||
2. **grm published** — latest grm tag exists in PyPI
|
|
||||||
3. **sso-bridge published** — latest sso_bridge tag exists in PyPI
|
|
||||||
4. **sso-bridge image built** — latest sso-bridge Docker image exists
|
|
||||||
5. **mattermost-oidc image built** — latest mattermost-oidc image exists
|
|
||||||
6. **infra pins match published versions** — no stale pins
|
|
||||||
7. **Nightly gate green** — `NIGHTLY_STATUS` is not `failed`
|
|
||||||
|
|
||||||
## Quick Check Commands
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Check latest devx version
|
|
||||||
curl -sS https://git.oblachno.oblachno.fyi/api/v1/repos/oblachno-oss/devx/releases/latest | python3 -c "import json,sys; print(json.load(sys.stdin).get('tag_name','?'))"
|
|
||||||
|
|
||||||
# Check pinned devx version in each repo
|
|
||||||
for repo in grm sso-bridge infra; do
|
|
||||||
echo -n "$repo: "; grep 'devx @' /home/emo/dev/ideas/oblachno/$repo/pyproject.toml | grep -oP 'v[\d.]+'
|
|
||||||
done
|
|
||||||
|
|
||||||
# Check latest sso-bridge image build
|
|
||||||
curl -sS -H "Authorization: token $DEVELOPER_GITEA_API_TOKEN" \
|
|
||||||
"https://git.oblachno.oblachno.fyi/api/v1/repos/oblachno/sso-bridge/actions/runs?per_page=5" \
|
|
||||||
| python3 -c "import json,sys; [print(r['id'],r['status'],r['conclusion']) for r in json.load(sys.stdin).get('workflow_runs',[]) if r.get('event')=='push']"
|
|
||||||
```
|
|
||||||
@@ -1,78 +0,0 @@
|
|||||||
# deployment-coordination
|
|
||||||
|
|
||||||
How grm releases propagate to infra. grm publishes a PyPI package and
|
|
||||||
auto-creates an infra dependency PR. Coordination ensures the PR is
|
|
||||||
merged before infra deploys.
|
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
Invoke this skill when:
|
|
||||||
- Changes to grm affect infra deployments
|
|
||||||
- Preparing a grm release that infra depends on
|
|
||||||
- Verifying infra has bumped to the latest grm version
|
|
||||||
- Coordinating a multi-repo change that includes grm
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- grm repo at `/home/emo/dev/ideas/oblachno/grm`
|
|
||||||
- `.env` with `DEVELOPER_GITEA_API_TOKEN`
|
|
||||||
- See `dependency-graph` skill for the full ecosystem map
|
|
||||||
|
|
||||||
## What grm Produces
|
|
||||||
|
|
||||||
grm publishes a Python package to the Gitea PyPI registry. infra pins it:
|
|
||||||
```toml
|
|
||||||
"grm @ git+https://git.oblachno.oblachno.fyi/oblachno-oss/grm.git@vX.Y.Z"
|
|
||||||
```
|
|
||||||
|
|
||||||
## Release Flow
|
|
||||||
|
|
||||||
1. PR merged to master
|
|
||||||
2. Post-merge workflow runs `devx.ci.release` — classifies changes
|
|
||||||
3. If user-facing changes: git-cliff bumps version, creates tag, pushes
|
|
||||||
4. `devx.ci.publish` builds and publishes to Gitea PyPI registry
|
|
||||||
5. `devx.ci.create_dependency_pr --repo oblachno/infra --package grm`
|
|
||||||
auto-creates an infra PR to bump the pinned grm version
|
|
||||||
|
|
||||||
## Downstream Consumer
|
|
||||||
|
|
||||||
| Repo | Pin location | Auto-bump? |
|
|
||||||
|------|-------------|------------|
|
|
||||||
| infra | `pyproject.toml` | Yes — auto PR created by post-merge |
|
|
||||||
|
|
||||||
## Coordinating a grm Change
|
|
||||||
|
|
||||||
1. **Merge grm PR** — wait for post-merge publish + dependency-PR creation
|
|
||||||
2. **Verify publish** — check the new tag:
|
|
||||||
```bash
|
|
||||||
curl -sS -H "Authorization: token $DEVELOPER_GITEA_API_TOKEN" \
|
|
||||||
https://git.oblachno.oblachno.fyi/api/v1/repos/oblachno-oss/grm/releases/latest \
|
|
||||||
| python3 -c "import json,sys; print(json.load(sys.stdin).get('tag_name','?'))"
|
|
||||||
```
|
|
||||||
3. **Find the auto-created infra PR** — check infra for open PRs with
|
|
||||||
`dependency` label or title containing `bump grm`
|
|
||||||
4. **Review and merge the infra dependency PR** — verify the version bump,
|
|
||||||
run `make pytest-cov` in infra, add `ready-to-merge`
|
|
||||||
5. **Verify infra staging deploy** — after infra merges, staging deploy
|
|
||||||
picks up the new grm version
|
|
||||||
|
|
||||||
## State Verification
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Current grm version
|
|
||||||
grep '__version__' /home/emo/dev/ideas/oblachno/grm/src/grm/__init__.py
|
|
||||||
|
|
||||||
# What infra pins
|
|
||||||
grep 'grm @' /home/emo/dev/ideas/oblachno/infra/pyproject.toml | grep -oP 'v[\d.]+'
|
|
||||||
|
|
||||||
# Check for open infra dependency PRs
|
|
||||||
curl -sS -H "Authorization: token $DEVELOPER_GITEA_API_TOKEN" \
|
|
||||||
"https://git.oblachno.oblachno.fyi/api/v1/repos/oblachno/infra/pulls?state=open" \
|
|
||||||
| python3 -c "import json,sys; [print(p['number'],p['title']) for p in json.load(sys.stdin) if 'grm' in p.get('title','').lower()]"
|
|
||||||
```
|
|
||||||
|
|
||||||
## Common Mistakes
|
|
||||||
|
|
||||||
- Merging the grm PR but ignoring the auto-created infra dependency PR
|
|
||||||
- Deploying infra before the dependency PR is merged — stale grm version
|
|
||||||
- Forgetting that grm also has an Ansible role used by infra at deploy time
|
|
||||||
@@ -2,28 +2,18 @@
|
|||||||
|
|
||||||
Quick reference for devx tools when working on this repo.
|
Quick reference for devx tools when working on this repo.
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
Invoke this skill when creating PRs, checking CI status, adding
|
|
||||||
labels, rebasing branches, or performing any PR lifecycle operation.
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- `.venv` exists (run `make setup` if not)
|
|
||||||
- `.env` with `DEVELOPER_GITEA_API_TOKEN`, `VIKUNJA_TOKEN`
|
|
||||||
|
|
||||||
## PR Workflow (use these, not raw git/tea/MCP)
|
## PR Workflow (use these, not raw git/tea/MCP)
|
||||||
|
|
||||||
| Task | Command |
|
| Task | Command |
|
||||||
|------|---------|
|
|------|---------|
|
||||||
| Create Vikunja task | `.venv/bin/python -m devx.tools.create_task --title "..." --description "..."` (make target doesn't forward args) |
|
| Create Vikunja task | `make create-task -- --title "..." --description "..."` |
|
||||||
| Create PR | `make create-pr` |
|
| Create PR | `make create-pr` |
|
||||||
| Push + create PR | `make push-with-pr` |
|
| Push + create PR | `make push-with-pr` |
|
||||||
| Check CI status | `make devx-pr-status` or `make devx-pr-status PR=42 WAIT=1` |
|
| Check CI status | `make devx-pr-status` or `make devx-pr-status PR=42 WAIT=1` |
|
||||||
| Fetch CI failure logs | `make devx-pr-logs` or `make devx-pr-logs PR=42 JOB=quality TAIL=50` |
|
| Fetch CI failure logs | `make devx-pr-logs` or `make devx-pr-logs PR=42 JOB=quality TAIL=50` |
|
||||||
| Add ready-to-merge label | `make devx-pr-label` or `make devx-pr-label PR=42` |
|
| Add ready-to-merge label | `make devx-pr-label` or `make devx-pr-label PR=42` |
|
||||||
| Rebase current branch | `make devx-rebase` |
|
| Rebase current branch | `make rebase` |
|
||||||
| Rebase PR via API | `make devx-pr-rebase` or `make pr-rebase PR=42` |
|
| Rebase PR via API | `make pr-rebase` or `make pr-rebase PR=42` |
|
||||||
|
|
||||||
## Auto-merge Behavior
|
## Auto-merge Behavior
|
||||||
|
|
||||||
|
|||||||
@@ -1,70 +0,0 @@
|
|||||||
# molecule-testing
|
|
||||||
|
|
||||||
Authoring and debugging `gitea_runner` molecule scenarios. For running
|
|
||||||
tests use the `testing-and-debugging` make targets — this covers
|
|
||||||
writing scenarios and fixing DIND/platform issues.
|
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
- Adding a molecule scenario for the `gitea_runner` role
|
|
||||||
- A scenario fails on platform setup, DIND, or registration mocking
|
|
||||||
- Reviewing scenario coverage for a role change
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- Docker running locally
|
|
||||||
- `.venv` exists (`make setup`)
|
|
||||||
|
|
||||||
## Scenario Layout
|
|
||||||
|
|
||||||
`ansible/roles/gitea_runner/molecule/<scenario>/`:
|
|
||||||
|
|
||||||
Current scenarios: `default`, `template-content`, `deregister`,
|
|
||||||
`multi-instance`, `update`, `remove`, `lifecycle`.
|
|
||||||
|
|
||||||
| File | Purpose |
|
|
||||||
|------|---------|
|
|
||||||
| `molecule.yml` | driver/platforms/provisioner config |
|
|
||||||
| `converge.yml` | applies the role |
|
|
||||||
| `verify.yml` | assertions scoped to the scenario |
|
|
||||||
| `prepare.yml` | optional host prep |
|
|
||||||
|
|
||||||
Scenario registration lives in `pyproject.toml` (scenario map used by
|
|
||||||
`devx.molecule` distribution in CI) — a new scenario MUST be
|
|
||||||
registered there or CI never runs it.
|
|
||||||
|
|
||||||
## molecule.yml Conventions
|
|
||||||
|
|
||||||
- Platform name/image/command are env-overridable via
|
|
||||||
`${MOLECULE_PLATFORM_*}` so all-platforms runs work.
|
|
||||||
- `remote_tmp: /tmp` in provisioner `config_options` — default temp
|
|
||||||
dir breaks in containers.
|
|
||||||
- `ANSIBLE_ROLES_PATH` must include the repo roles root.
|
|
||||||
- Use `inventory.group_vars` to isolate the scenario: disable
|
|
||||||
unrelated features rather than editing tasks.
|
|
||||||
- Runner registration in tests is mocked/faked — scenarios must not
|
|
||||||
require a live Gitea instance; check how existing scenarios stub
|
|
||||||
the registration/token flow before adding API calls.
|
|
||||||
|
|
||||||
## Debugging
|
|
||||||
|
|
||||||
```bash
|
|
||||||
cd ansible/roles/gitea_runner
|
|
||||||
molecule test -s <scenario>
|
|
||||||
molecule converge -s <scenario>
|
|
||||||
molecule login -s <scenario>
|
|
||||||
```
|
|
||||||
|
|
||||||
- "Failed to create temporary directory" → `remote_tmp: /tmp` missing.
|
|
||||||
- Idempotence failures → find the changed task on second converge.
|
|
||||||
- Registration/API timeouts → the scenario hit a real endpoint —
|
|
||||||
stub it like the existing scenarios do.
|
|
||||||
|
|
||||||
## Common Mistakes
|
|
||||||
|
|
||||||
- Adding a scenario without registering it in `pyproject.toml` —
|
|
||||||
silently untested.
|
|
||||||
- Hardcoding the platform image — keep `${MOLECULE_PLATFORM_*}`
|
|
||||||
overrides.
|
|
||||||
- Calling the real Gitea API in converge — scenarios must be
|
|
||||||
self-contained; mock the registration path.
|
|
||||||
@@ -1,73 +0,0 @@
|
|||||||
# runner-ops
|
|
||||||
|
|
||||||
Operating the Gitea Actions runner fleet: registration lifecycle,
|
|
||||||
stale-runner cleanup, image pruning, and safe debugging. Core code:
|
|
||||||
`src/grm/runner_manager.py`, `src/grm/executor.py`,
|
|
||||||
`src/grm/registry.py`.
|
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
- Runners go offline, stall, or pile up stale registrations
|
|
||||||
- Runner hosts need install/update/remove/deregister operations
|
|
||||||
- Disk pressure on runner hosts (image/container accumulation)
|
|
||||||
- Working on S08 (leases, physical-host admission, disk watermarks)
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- `.env` with Gitea admin token for API operations
|
|
||||||
- SSH access to runner hosts for Ansible-driven lifecycle
|
|
||||||
- Runner registrations visible via admin API:
|
|
||||||
`GET /api/v1/admin/actions/runners`
|
|
||||||
|
|
||||||
## Architecture
|
|
||||||
|
|
||||||
- `RunnerManager` orchestrates install/update/lifecycle via
|
|
||||||
`AnsibleExecutor` against the `gitea_runner` role; `RunnerRegistry`
|
|
||||||
tracks local runner state.
|
|
||||||
- Runners execute jobs in Docker (`docker` label) — every job gets a
|
|
||||||
fresh container from `ci-base`/`ci-quality`/`ci-full` images.
|
|
||||||
- Molecule jobs nest containers (DIND) — privileged, `SYS_ADMIN`,
|
|
||||||
`/var/lib/docker` volume.
|
|
||||||
|
|
||||||
## Lifecycle Operations
|
|
||||||
|
|
||||||
| Task | Entry point |
|
|
||||||
|------|-------------|
|
|
||||||
| Install/update runners | `grm` CLI → `RunnerManager` (Ansible) |
|
|
||||||
| Stale registration cleanup | `scripts/cleanup_stale_runners.py` — deletes runners offline >1h via `DELETE /api/v1/admin/actions/runners/{id}` |
|
|
||||||
| Image pruning | `scripts/prune_runner_images.py` — reclaims disk from old CI image versions |
|
|
||||||
|
|
||||||
Stale registrations accumulate when a host is rebuilt, re-registered,
|
|
||||||
or its runner process dies unrecoverably — clean them before capacity
|
|
||||||
accounting.
|
|
||||||
|
|
||||||
## Debugging a Stuck Runner
|
|
||||||
|
|
||||||
1. Check registration state via admin API (offline vs online).
|
|
||||||
2. SSH to the host: `systemctl status` the runner service / inspect
|
|
||||||
`docker ps` for orphaned job containers.
|
|
||||||
3. Orphaned molecule containers: safe to remove ONLY when no molecule
|
|
||||||
run is active — check runner logs first (`runner-ops` counterpart
|
|
||||||
of "don't force-remove active containers", fixed in GRM-166/167).
|
|
||||||
4. Disk pressure: check `/var/lib/docker` usage, then
|
|
||||||
`prune_runner_images.py` — never blanket `docker system prune`
|
|
||||||
while jobs may be mid-flight.
|
|
||||||
|
|
||||||
## S08-Relevant Rules
|
|
||||||
|
|
||||||
- Runner admission must be per physical host — a runner that shares
|
|
||||||
hardware must declare capacity, not just labels.
|
|
||||||
- Cleanup must never remove a container a live job owns — ownership
|
|
||||||
check before any force-removal.
|
|
||||||
- Disk watermark logic belongs in the role/scripts, not ad-hoc
|
|
||||||
cron `docker prune`.
|
|
||||||
|
|
||||||
## Common Mistakes
|
|
||||||
|
|
||||||
- `docker system prune -a` on a runner host — kills in-flight job
|
|
||||||
containers and image cache mid-run.
|
|
||||||
- Deleting an offline runner registration while the host still runs
|
|
||||||
the service — it re-registers and duplicates; stop the service
|
|
||||||
first.
|
|
||||||
- Treating molecule DIND containers as junk — they belong to an
|
|
||||||
active scenario; check timestamps and runner logs.
|
|
||||||
@@ -1,142 +0,0 @@
|
|||||||
# skill-creation
|
|
||||||
|
|
||||||
How to create, validate, and maintain Devin skills. Skills must be
|
|
||||||
clear, succinct, and actionable — no AI slop.
|
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
Invoke this skill when:
|
|
||||||
- Creating a new skill
|
|
||||||
- Amending an existing skill
|
|
||||||
- Evaluating whether a skill is needed
|
|
||||||
- Reviewing a PR that adds or modifies skills
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- Skill directory: `.devin/skills/<skill-name>/SKILL.md`
|
|
||||||
- Validator: `tests/test_skills.py` (infra) or reference to it
|
|
||||||
- Tests: `tests/unit/test_skills_validation.py` (infra)
|
|
||||||
|
|
||||||
## When to Create a Skill
|
|
||||||
|
|
||||||
Create a skill when:
|
|
||||||
- An agent struggles with a task repeatedly (branch hygiene, PR order)
|
|
||||||
- A workflow has non-obvious ordering constraints (deployment coordination)
|
|
||||||
- A task requires specific tool usage over raw commands (CI monitoring)
|
|
||||||
- Multiple agents need shared context (dependency graph)
|
|
||||||
|
|
||||||
Do NOT create a skill for:
|
|
||||||
- One-off tasks (use a spec instead)
|
|
||||||
- Tasks already covered by AGENTS.md
|
|
||||||
- Tasks that are obvious from the Makefile or README
|
|
||||||
- Tasks that change frequently (skills should be stable)
|
|
||||||
|
|
||||||
## Skill Structure
|
|
||||||
|
|
||||||
Every skill MUST have:
|
|
||||||
|
|
||||||
```markdown
|
|
||||||
# <skill-name>
|
|
||||||
|
|
||||||
One-line description of what the skill does.
|
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
2-4 bullet points describing when to use this skill.
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
What must exist before using the skill (venv, .env, tools).
|
|
||||||
|
|
||||||
## <Core Content>
|
|
||||||
|
|
||||||
The actual guidance. Keep it actionable.
|
|
||||||
|
|
||||||
## Verification (if applicable)
|
|
||||||
|
|
||||||
How to verify the skill's guidance works.
|
|
||||||
|
|
||||||
## Common Mistakes (if applicable)
|
|
||||||
|
|
||||||
What agents get wrong without this skill.
|
|
||||||
```
|
|
||||||
|
|
||||||
## Quality Standards
|
|
||||||
|
|
||||||
### Do
|
|
||||||
- **Be specific.** Reference exact make targets, file paths, commands.
|
|
||||||
- **Be concise.** Each section should be scannable in under 30 seconds.
|
|
||||||
- **Be actionable.** Every paragraph should tell the agent what to DO.
|
|
||||||
- **Use tables** for command reference, mappings, and comparisons.
|
|
||||||
- **Use code blocks** for commands the agent should run.
|
|
||||||
- **Link to other skills** when related (e.g., "See `dependency-graph` skill").
|
|
||||||
|
|
||||||
### Don't
|
|
||||||
- **No preamble.** Don't start with "This skill helps agents..." — just state what it does.
|
|
||||||
- **No filler.** Don't repeat information from AGENTS.md or other skills.
|
|
||||||
- **No vague advice.** "Be careful with branches" is useless. "Run `git branch --show-current` before every commit" is useful.
|
|
||||||
- **No AI slop.** Don't write "In this comprehensive guide, we will explore..." — just give the guidance.
|
|
||||||
- **No redundant sections.** If "Common Mistakes" would repeat "When to Invoke", skip it.
|
|
||||||
- **No marketing.** Don't describe the skill as "powerful" or "comprehensive".
|
|
||||||
|
|
||||||
## Scope Rules
|
|
||||||
|
|
||||||
- **One skill per concern.** Don't mix branch hygiene with CI monitoring.
|
|
||||||
- **Project-specific, not generic.** Skills reference this repo's make targets, file paths, and conventions — not abstract advice.
|
|
||||||
- **Shared skills must be identical across repos.** Use `SHARED_SKILLS` in the validator to enforce this.
|
|
||||||
- **Per-repo skills must reflect that repo's reality.** Don't copy infra-specific targets to sso-bridge.
|
|
||||||
|
|
||||||
## Automated Validation
|
|
||||||
|
|
||||||
Every skill must pass the validator (`tests/test_skills.py`). The validator checks:
|
|
||||||
|
|
||||||
1. **Structure** — H1 title, "When to Invoke" section, "Prerequisites" section
|
|
||||||
2. **Commands** — referenced `make <target>` commands exist in Makefile or devx.mak
|
|
||||||
3. **Paths** — referenced file paths exist in the repo
|
|
||||||
4. **Shared skills** — identical content across repos (SHA-256 comparison)
|
|
||||||
5. **No drift** — no references to nonexistent commands or files
|
|
||||||
|
|
||||||
Run the validator:
|
|
||||||
```bash
|
|
||||||
python3 tests/test_skills.py --repo infra --repo sso-bridge
|
|
||||||
```
|
|
||||||
|
|
||||||
## Effectiveness Evaluation
|
|
||||||
|
|
||||||
### Static Checks (automated, CI)
|
|
||||||
|
|
||||||
The validator runs in CI as part of `make pytest-cov`. A failing skill
|
|
||||||
test blocks the PR. This catches:
|
|
||||||
- Missing sections
|
|
||||||
- Invalid commands
|
|
||||||
- Broken file references
|
|
||||||
- Cross-repo drift
|
|
||||||
|
|
||||||
### Runtime Metrics (manual, periodic)
|
|
||||||
|
|
||||||
Track these signals to evaluate skill effectiveness:
|
|
||||||
- **Skill invocation frequency** — how often agents invoke the skill
|
|
||||||
- **Success rate when invoked** — did the skill prevent the mistake it targets?
|
|
||||||
- **Feedback issues** — agents create Gitea issues with `feedback` label when a skill is unclear or wrong
|
|
||||||
- **Mistake recurrence** — if agents still make the mistake the skill targets, the skill needs improvement
|
|
||||||
|
|
||||||
### Retrospective Review
|
|
||||||
|
|
||||||
Periodically (monthly or after major incidents) review skills:
|
|
||||||
1. List all skills and their last-modified dates
|
|
||||||
2. Check for feedback issues tagged `skill-improvement`
|
|
||||||
3. Verify referenced commands still exist (run validator)
|
|
||||||
4. Remove skills that are no longer relevant
|
|
||||||
5. Update skills where mistakes still recur
|
|
||||||
6. Document lessons in this skill's "Common Mistakes" section
|
|
||||||
|
|
||||||
## Creating a New Skill — Checklist
|
|
||||||
|
|
||||||
- [ ] Identify the repeated struggle or non-obvious workflow
|
|
||||||
- [ ] Check no existing skill covers it
|
|
||||||
- [ ] Write the skill following the structure above
|
|
||||||
- [ ] Run `python3 tests/test_skills.py` — must pass
|
|
||||||
- [ ] Run `make pytest-cov` — must pass with 100% coverage
|
|
||||||
- [ ] If shared across repos, copy identical content to each repo
|
|
||||||
- [ ] Add the skill to `SHARED_SKILLS` in the validator if shared
|
|
||||||
- [ ] Create PR, verify CI passes, merge
|
|
||||||
@@ -1,14 +1,5 @@
|
|||||||
# Spec-Driven Development
|
# Spec-Driven Development
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
Invoke this skill when starting any change — every PR requires a spec
|
|
||||||
at `docs/specs/<TASK-ID>.md` that CI validates before merge.
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- A Vikunja task ID (`GRM-N`) — see `vikunja-tasks` skill
|
|
||||||
|
|
||||||
## Overview
|
## Overview
|
||||||
|
|
||||||
Every change starts with a spec. No spec, no code. No code, no PR.
|
Every change starts with a spec. No spec, no code. No code, no PR.
|
||||||
|
|||||||
@@ -3,17 +3,6 @@
|
|||||||
Make targets for testing, debugging, and CI investigation. **Use these
|
Make targets for testing, debugging, and CI investigation. **Use these
|
||||||
instead of raw `pytest`, `ruff`, or `molecule` commands.**
|
instead of raw `pytest`, `ruff`, or `molecule` commands.**
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
Invoke this skill when running tests, investigating CI failures, or
|
|
||||||
debugging molecule scenarios. Also invoke when asked to "run tests",
|
|
||||||
"check coverage", or "debug a failure".
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- `.venv` exists (run `make setup` if not)
|
|
||||||
- For molecule tests: Docker is running
|
|
||||||
|
|
||||||
## Why Make Targets
|
## Why Make Targets
|
||||||
|
|
||||||
Make targets encapsulate the correct venv activation, PYTHONPATH, env
|
Make targets encapsulate the correct venv activation, PYTHONPATH, env
|
||||||
@@ -42,8 +31,9 @@ produces false failures (missing dependencies, wrong Python version).
|
|||||||
|
|
||||||
| Task | Command | Notes |
|
| Task | Command | Notes |
|
||||||
|------|---------|-------|
|
|------|---------|-------|
|
||||||
| All scenarios | `make molecule` | All 7 scenarios on Ubuntu 22.04 |
|
| All scenarios | `make molecule` | All 6 scenarios on Ubuntu 22.04 |
|
||||||
| All platforms | `make molecule-all` | All 7 scenarios on all 4 OSes |
|
| All platforms | `make molecule-all` | All 6 scenarios on all 4 OSes |
|
||||||
|
| Parallel | `make molecule-all-parallel` | MOLECULE_JOBS=4 |
|
||||||
|
|
||||||
### Spec-Driven Workflow
|
### Spec-Driven Workflow
|
||||||
|
|
||||||
@@ -56,12 +46,12 @@ CI validates the spec before running expensive jobs.
|
|||||||
**Before pushing any branch:**
|
**Before pushing any branch:**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
make lint-all && make pytest-cov
|
make pre-push
|
||||||
```
|
```
|
||||||
|
|
||||||
This runs all linters + unit tests with coverage. The pre-push git
|
This runs `lint-all` + `pytest-cov`. The pre-push git hook only
|
||||||
hook only validates the Vikunja task exists — it does NOT run tests.
|
validates the Vikunja task exists — it does NOT run tests. You must
|
||||||
Run the checks manually (there is no `pre-push` target here).
|
run `make pre-push` manually.
|
||||||
|
|
||||||
## CI Failure Investigation
|
## CI Failure Investigation
|
||||||
|
|
||||||
@@ -69,7 +59,7 @@ When investigating a CI failure:
|
|||||||
|
|
||||||
1. **Fetch logs via MCP** — use `mcp_call_tool` with gitea server,
|
1. **Fetch logs via MCP** — use `mcp_call_tool` with gitea server,
|
||||||
`actions_run_read` method, `download_job_log` tool
|
`actions_run_read` method, `download_job_log` tool
|
||||||
2. **Reproduce locally** — use `make pytest-cov` or `make lint-all`
|
2. **Reproduce locally** — use `make pytest-cov` or `make lint-ci`
|
||||||
depending on which CI job failed
|
depending on which CI job failed
|
||||||
3. **Never run raw pytest** — always use the make target
|
3. **Never run raw pytest** — always use the make target
|
||||||
|
|
||||||
|
|||||||
@@ -1,74 +0,0 @@
|
|||||||
# vikunja-tasks
|
|
||||||
|
|
||||||
Vikunja task lifecycle beyond `create`: querying status, closing, and
|
|
||||||
recovering when the tracker is unreachable.
|
|
||||||
|
|
||||||
## When to Invoke
|
|
||||||
|
|
||||||
- Creating, closing, or checking a Vikunja task
|
|
||||||
- A spec workflow step needs the task ID or done state
|
|
||||||
- `vikunja.oblachno.oblachno.fyi` fails to resolve / times out
|
|
||||||
|
|
||||||
## Prerequisites
|
|
||||||
|
|
||||||
- `.env` with `VIKUNJA_TOKEN`
|
|
||||||
- Project ID comes from `[tool.devx]` in `pyproject.toml`
|
|
||||||
(`DEVX_VIKUNJA_PROJECT_ID`)
|
|
||||||
|
|
||||||
## Create
|
|
||||||
|
|
||||||
`make create-task` does **not** forward arguments — call the module:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
.venv/bin/python -m devx.tools.create_task \
|
|
||||||
--title "Task title (no GRM-N prefix)" \
|
|
||||||
--description "<h2>Context</h2><p>...</p>"
|
|
||||||
```
|
|
||||||
|
|
||||||
Prints `GRM-N` + next steps. Title must not include the task-ID
|
|
||||||
prefix (auto-merge prepends it; a manual prefix double-prefixes the
|
|
||||||
PR title and fails validation).
|
|
||||||
|
|
||||||
## Query / Close
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# Task details (ID = numeric part of GRM-N)
|
|
||||||
curl -sf -H "Authorization: Bearer $VIKUNJA_TOKEN" \
|
|
||||||
"https://vikunja.oblachno.oblachno.fyi/api/v1/tasks/<N>"
|
|
||||||
|
|
||||||
# Close: mark done
|
|
||||||
curl -sf -X POST -H "Authorization: Bearer $VIKUNJA_TOKEN" \
|
|
||||||
-H "Content-Type: application/json" -d '{"done":true}' \
|
|
||||||
"https://vikunja.oblachno.oblachno.fyi/api/v1/tasks/<N>"
|
|
||||||
```
|
|
||||||
|
|
||||||
Post-merge automation marks the task done when the PR squash-merges —
|
|
||||||
manual close is only needed for abandoned/superseded tasks.
|
|
||||||
|
|
||||||
## Task-ID / Spec Collisions
|
|
||||||
|
|
||||||
Vikunja IDs can collide with historical spec files (an old task reused
|
|
||||||
the number). Convention: preserve the old file as
|
|
||||||
`docs/specs/<ID>-<topic>-historical.md`, then write the new spec at
|
|
||||||
`docs/specs/<ID>.md`. Check `git log` on the existing spec before
|
|
||||||
moving it.
|
|
||||||
|
|
||||||
## Tracker Unreachable
|
|
||||||
|
|
||||||
If the Vikunja host fails DNS/TLS:
|
|
||||||
|
|
||||||
1. Don't block the whole workflow — record the intended task title in
|
|
||||||
the spec draft and retry `create_task` before branching.
|
|
||||||
2. Never invent an ID — branch/PR titles must match a real task or
|
|
||||||
`pre_push_check` / auto-merge validation fails.
|
|
||||||
3. DNS failures observed so far were transient; retry after a few
|
|
||||||
minutes before escalating.
|
|
||||||
|
|
||||||
## Common Mistakes
|
|
||||||
|
|
||||||
- `make create-task -- --title ...` — args are dropped; use the module
|
|
||||||
call above (forwarding fix is S11 scope).
|
|
||||||
- Including `GRM-N:` in the task title — double prefix breaks
|
|
||||||
auto-merge.
|
|
||||||
- Closing a task whose PR is still open — auto-merge's post-merge
|
|
||||||
step handles the close; manual close confuses the audit trail.
|
|
||||||
@@ -283,7 +283,7 @@ jobs:
|
|||||||
needs.validate.result == 'success'
|
needs.validate.result == 'success'
|
||||||
runs-on: docker
|
runs-on: docker
|
||||||
container: git.oblachno.oblachno.fyi/oblachno-oss/runner-images/ci-base:latest
|
container: git.oblachno.oblachno.fyi/oblachno-oss/runner-images/ci-base:latest
|
||||||
timeout-minutes: 50
|
timeout-minutes: 10
|
||||||
defaults:
|
defaults:
|
||||||
run:
|
run:
|
||||||
shell: bash
|
shell: bash
|
||||||
@@ -299,18 +299,16 @@ jobs:
|
|||||||
run: make setup-image EXTRAS=ci
|
run: make setup-image EXTRAS=ci
|
||||||
- name: Post approval review
|
- name: Post approval review
|
||||||
env:
|
env:
|
||||||
DEVELOPER_GITEA_API_TOKEN: ${{ secrets.DEVELOPER_GITEA_API_TOKEN }}
|
REVIEWER_GITEA_API_TOKEN: ${{ secrets.REVIEWER_GITEA_API_TOKEN }}
|
||||||
PR_NUMBER: ${{ github.event.number }}
|
PR_NUMBER: ${{ github.event.number }}
|
||||||
GITHUB_SERVER_URL: ${{ github.server_url }}
|
GITHUB_SERVER_URL: ${{ github.server_url }}
|
||||||
GITHUB_REPOSITORY: ${{ github.repository }}
|
GITHUB_REPOSITORY: ${{ github.repository }}
|
||||||
run: |
|
run: |
|
||||||
. .venv/bin/activate 2>/dev/null || true
|
. .venv/bin/activate 2>/dev/null || true
|
||||||
# Post APPROVE review via Gitea API to satisfy branch protection
|
# Post APPROVE review via Gitea API to satisfy branch protection
|
||||||
# Uses DEVELOPER_GITEA_API_TOKEN (kireto) — a different user than
|
|
||||||
# the PR creator — so Gitea counts the approval (no self-approvals).
|
|
||||||
curl -s -X POST \
|
curl -s -X POST \
|
||||||
"${GITHUB_SERVER_URL}/api/v1/repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}/reviews" \
|
"${GITHUB_SERVER_URL}/api/v1/repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}/reviews" \
|
||||||
-H "Authorization: token ${DEVELOPER_GITEA_API_TOKEN}" \
|
-H "Authorization: token ${REVIEWER_GITEA_API_TOKEN}" \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{"event":"APPROVED","body":"Auto-approved: all CI checks passed (validate, molecule-tests)."}' \
|
-d '{"event":"APPROVED","body":"Auto-approved: all CI checks passed (validate, molecule-tests)."}' \
|
||||||
|| echo "::warning::Failed to post approval review (best-effort)."
|
|| echo "::warning::Failed to post approval review (best-effort)."
|
||||||
@@ -321,7 +319,7 @@ jobs:
|
|||||||
run: |
|
run: |
|
||||||
. .venv/bin/activate 2>/dev/null || true
|
. .venv/bin/activate 2>/dev/null || true
|
||||||
# Poll commit status until all required checks pass or fail
|
# Poll commit status until all required checks pass or fail
|
||||||
MAX_WAIT=2400 # 40 minutes — covers the ~25-min molecule suite
|
MAX_WAIT=600 # 10 minutes
|
||||||
ELAPSED=0
|
ELAPSED=0
|
||||||
while [ $ELAPSED -lt $MAX_WAIT ]; do
|
while [ $ELAPSED -lt $MAX_WAIT ]; do
|
||||||
STATUS=$(curl -s -H "Authorization: token $CI_GITEA_API_TOKEN" \
|
STATUS=$(curl -s -H "Authorization: token $CI_GITEA_API_TOKEN" \
|
||||||
|
|||||||
@@ -111,7 +111,7 @@ The auto-merge workflow enforces the APPROVE review check programmatically
|
|||||||
as a defense-in-depth measure, but branch protection is the primary gate.
|
as a defense-in-depth measure, but branch protection is the primary gate.
|
||||||
|
|
||||||
### 1. Create Vikunja Task
|
### 1. Create Vikunja Task
|
||||||
Create a task in Vikunja project 6 via `.venv/bin/python -m devx.tools.create_task --title "Task title" --description "<h2>...</h2>"` (make target does not forward args) (requires `VIKUNJA_TOKEN` in `.env`). This prints the `GRM-N` identifier and next-step instructions.
|
Create a task in Vikunja project 6 via `make create-task -- --title "Task title" --description "<h2>...</h2>"` (requires `VIKUNJA_TOKEN` in `.env`). This prints the `GRM-N` identifier and next-step instructions.
|
||||||
|
|
||||||
**IMPORTANT:** The task title must NOT include the `GRM-N:` prefix.
|
**IMPORTANT:** The task title must NOT include the `GRM-N:` prefix.
|
||||||
The `make create-pr` and `check_auto_merge_ready` commands automatically
|
The `make create-pr` and `check_auto_merge_ready` commands automatically
|
||||||
|
|||||||
@@ -2,37 +2,6 @@
|
|||||||
|
|
||||||
All notable changes to this project will be documented in this file.
|
All notable changes to this project will be documented in this file.
|
||||||
|
|
||||||
## [0.23.3] - 2026-09-18
|
|
||||||
|
|
||||||
### Bug Fixes
|
|
||||||
|
|
||||||
- Raise auto-merge molecule wait to cover suite duration
|
|
||||||
- Harden stall-detection enumeration and timestamp parsing
|
|
||||||
|
|
||||||
## [0.23.2] - 2026-09-15
|
|
||||||
|
|
||||||
### Bug Fixes
|
|
||||||
|
|
||||||
- Restrict healthcheck disk cleanup to exited/dead containers
|
|
||||||
|
|
||||||
## [0.23.1] - 2026-09-15
|
|
||||||
|
|
||||||
### Bug Fixes
|
|
||||||
|
|
||||||
- Restrict docker-prune to exited containers
|
|
||||||
|
|
||||||
## [0.23.0] - 2026-08-28
|
|
||||||
|
|
||||||
### Features
|
|
||||||
|
|
||||||
- Add pre-cache timer, force_pull, and Docker socket options to runner config
|
|
||||||
|
|
||||||
## [0.22.1] - 2026-08-26
|
|
||||||
|
|
||||||
### Bug Fixes
|
|
||||||
|
|
||||||
- Use kireto token for auto-merge approval review
|
|
||||||
|
|
||||||
## [0.22.0] - 2026-08-25
|
## [0.22.0] - 2026-08-25
|
||||||
|
|
||||||
### Features
|
### Features
|
||||||
|
|||||||
@@ -8,12 +8,12 @@ Each runner runs in an isolated **rootless Docker** environment under a dedicate
|
|||||||
|
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases)
|
||||||
[](https://www.python.org/downloads/)
|
[](https://www.python.org/downloads/)
|
||||||
|
|
||||||
## Why GRM?
|
## Why GRM?
|
||||||
|
|
||||||
|
|||||||
@@ -50,11 +50,6 @@ gitea_runner_healthcheck_disk_threshold: 70
|
|||||||
gitea_runner_healthcheck_disk_critical: 75
|
gitea_runner_healthcheck_disk_critical: 75
|
||||||
gitea_runner_healthcheck_script_path: "{{ gitea_runner_config_dir }}/healthcheck.sh"
|
gitea_runner_healthcheck_script_path: "{{ gitea_runner_config_dir }}/healthcheck.sh"
|
||||||
|
|
||||||
# CI job containers older than this many minutes get an exec-responsiveness
|
|
||||||
# probe; a timeout writes one diagnostics bundle per container for
|
|
||||||
# post-mortem analysis of recurring ~20min exec/archive stalls (GRM-168).
|
|
||||||
gitea_runner_stall_minutes: 15
|
|
||||||
|
|
||||||
# Auto-recovery: when the healthcheck detects an unregistered runner, it
|
# Auto-recovery: when the healthcheck detects an unregistered runner, it
|
||||||
# can automatically re-register if a Gitea API token is provided.
|
# can automatically re-register if a Gitea API token is provided.
|
||||||
# The token needs admin or org-level access to fetch registration tokens.
|
# The token needs admin or org-level access to fetch registration tokens.
|
||||||
@@ -95,22 +90,6 @@ gitea_runner_remove_user: true
|
|||||||
gitea_runner_log_level: "info"
|
gitea_runner_log_level: "info"
|
||||||
gitea_runner_container_label: "gitea-runner=true"
|
gitea_runner_container_label: "gitea-runner=true"
|
||||||
gitea_runner_file: ".runner"
|
gitea_runner_file: ".runner"
|
||||||
# force_pull: when false (default), the runner reuses locally cached images
|
|
||||||
# instead of pulling on every job. Pre-cached images (via the pre-cache timer
|
|
||||||
# or pre_pull_images task) eliminate registry thundering-herd when all runners
|
|
||||||
# start jobs simultaneously.
|
|
||||||
gitea_runner_force_pull: false
|
|
||||||
# Container options passed to `docker run` for CI job containers.
|
|
||||||
# Mounts the host rootless Docker socket as /run/host-docker.sock so
|
|
||||||
# start_docker.py inside the container can detect and use the host daemon
|
|
||||||
# (full disk, no nested DinD) instead of starting an inner dockerd.
|
|
||||||
gitea_runner_container_options: "-v /run/user/{{ gitea_runner_uid }}/docker.sock:/run/host-docker.sock"
|
|
||||||
# Volumes allowed in CI job containers (validated by the runner against
|
|
||||||
# container.options and job-level volumes). Must include the host Docker
|
|
||||||
# socket mount target.
|
|
||||||
gitea_runner_valid_volumes:
|
|
||||||
- "/run/host-docker.sock"
|
|
||||||
- "/run/user/{{ gitea_runner_uid }}/docker.sock"
|
|
||||||
|
|
||||||
# Containerd version pinning — Docker 28.x vendors containerd v2.1.x internally.
|
# Containerd version pinning — Docker 28.x vendors containerd v2.1.x internally.
|
||||||
# containerd.io >= 2.3 ships a shim that returns a protobuf BootstrapResult which
|
# containerd.io >= 2.3 ships a shim that returns a protobuf BootstrapResult which
|
||||||
@@ -156,11 +135,3 @@ gitea_runner_docker_ipv6_cidr: "fd00:dead:beef::/48"
|
|||||||
# disk-space prune only removes dangling images, so pre-pulled tagged images persist.
|
# disk-space prune only removes dangling images, so pre-pulled tagged images persist.
|
||||||
# Set to [] to skip pre-pulling. Images are pulled as the runner user via rootless Docker.
|
# Set to [] to skip pre-pulling. Images are pulled as the runner user via rootless Docker.
|
||||||
gitea_runner_pre_pull_images: []
|
gitea_runner_pre_pull_images: []
|
||||||
|
|
||||||
# Pre-cache timer: periodically pulls the runner container image so it stays
|
|
||||||
# fresh in the local Docker cache. This prevents thundering-herd registry
|
|
||||||
# timeouts when all runners start CI jobs simultaneously with empty caches.
|
|
||||||
# Runs every 6 hours (aligned with prune schedule). Set to empty string to
|
|
||||||
# disable the timer.
|
|
||||||
gitea_runner_pre_cache_schedule: "*-*-* 00/6:30:00"
|
|
||||||
gitea_runner_pre_cache_images: "{{ gitea_runner_pre_pull_images }}"
|
|
||||||
|
|||||||
@@ -48,7 +48,6 @@
|
|||||||
that:
|
that:
|
||||||
- "'Type=oneshot' in prune_service.content | b64decode"
|
- "'Type=oneshot' in prune_service.content | b64decode"
|
||||||
- "'docker rm -f' in prune_service.content | b64decode"
|
- "'docker rm -f' in prune_service.content | b64decode"
|
||||||
- "'status=exited' in prune_service.content | b64decode"
|
|
||||||
- "'GITEA-ACTIONS-TASK' in prune_service.content | b64decode"
|
- "'GITEA-ACTIONS-TASK' in prune_service.content | b64decode"
|
||||||
- "'docker system prune -af' in prune_service.content | b64decode"
|
- "'docker system prune -af' in prune_service.content | b64decode"
|
||||||
- "'docker network prune' in prune_service.content | b64decode"
|
- "'docker network prune' in prune_service.content | b64decode"
|
||||||
@@ -112,8 +111,6 @@
|
|||||||
- "'docker network prune' in healthcheck_script.content | b64decode"
|
- "'docker network prune' in healthcheck_script.content | b64decode"
|
||||||
- "'status=removing' in healthcheck_script.content | b64decode"
|
- "'status=removing' in healthcheck_script.content | b64decode"
|
||||||
- "'status=stopping' in healthcheck_script.content | b64decode"
|
- "'status=stopping' in healthcheck_script.content | b64decode"
|
||||||
- "'status=exited' in healthcheck_script.content | b64decode"
|
|
||||||
- "'status=dead' in healthcheck_script.content | b64decode"
|
|
||||||
- "gitea_runner_healthcheck_disk_threshold | string in healthcheck_script.content | b64decode"
|
- "gitea_runner_healthcheck_disk_threshold | string in healthcheck_script.content | b64decode"
|
||||||
- "gitea_runner_healthcheck_disk_critical | string in healthcheck_script.content | b64decode"
|
- "gitea_runner_healthcheck_disk_critical | string in healthcheck_script.content | b64decode"
|
||||||
fail_msg: "Healthcheck script template is missing expected content"
|
fail_msg: "Healthcheck script template is missing expected content"
|
||||||
|
|||||||
@@ -20,9 +20,6 @@
|
|||||||
- name: Include pre-pull images
|
- name: Include pre-pull images
|
||||||
ansible.builtin.include_tasks: pre_pull_images.yml
|
ansible.builtin.include_tasks: pre_pull_images.yml
|
||||||
|
|
||||||
- name: Include pre-cache timer
|
|
||||||
ansible.builtin.include_tasks: pre_cache.yml
|
|
||||||
|
|
||||||
- name: Include integration test
|
- name: Include integration test
|
||||||
ansible.builtin.include_tasks: integration_test.yml
|
ansible.builtin.include_tasks: integration_test.yml
|
||||||
when: not gitea_runner_skip_registration
|
when: not gitea_runner_skip_registration
|
||||||
|
|||||||
@@ -1,67 +0,0 @@
|
|||||||
---
|
|
||||||
# Periodic timer that pre-pulls CI runner images into the local Docker cache.
|
|
||||||
# Prevents thundering-herd registry timeouts when all runners start jobs
|
|
||||||
# simultaneously with empty/stale caches. Runs every 6 hours (configurable).
|
|
||||||
# The prune timer removes dangling images but NOT tagged ones, so pre-pulled
|
|
||||||
# images persist between runs.
|
|
||||||
|
|
||||||
- name: Create docker-pull-images user service file
|
|
||||||
ansible.builtin.template:
|
|
||||||
src: docker-pull-images.service.j2
|
|
||||||
dest: "{{ gitea_runner_home }}/.config/systemd/user/docker-pull-images.service"
|
|
||||||
owner: "{{ gitea_runner_service_user }}"
|
|
||||||
group: "{{ gitea_runner_service_user }}"
|
|
||||||
mode: "0644"
|
|
||||||
register: gitea_runner_pre_cache_service
|
|
||||||
|
|
||||||
- name: Create docker-pull-images user timer file
|
|
||||||
ansible.builtin.template:
|
|
||||||
src: docker-pull-images.timer.j2
|
|
||||||
dest: "{{ gitea_runner_home }}/.config/systemd/user/docker-pull-images.timer"
|
|
||||||
owner: "{{ gitea_runner_service_user }}"
|
|
||||||
group: "{{ gitea_runner_service_user }}"
|
|
||||||
mode: "0644"
|
|
||||||
register: gitea_runner_pre_cache_timer
|
|
||||||
|
|
||||||
- name: Reload systemd user daemon for pre-cache timer
|
|
||||||
ansible.builtin.command: systemctl --user daemon-reload
|
|
||||||
become: true
|
|
||||||
become_user: "{{ gitea_runner_service_user }}"
|
|
||||||
environment:
|
|
||||||
XDG_RUNTIME_DIR: "/run/user/{{ gitea_runner_uid }}"
|
|
||||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ gitea_runner_uid | default(0) }}/bus"
|
|
||||||
changed_when: true
|
|
||||||
when:
|
|
||||||
- gitea_runner_systemd_available.stat.exists
|
|
||||||
- gitea_runner_docker_rootless_setup
|
|
||||||
- gitea_runner_pre_cache_service is changed or gitea_runner_pre_cache_timer is changed
|
|
||||||
- gitea_runner_pre_cache_schedule | length > 0
|
|
||||||
- gitea_runner_pre_cache_images | length > 0
|
|
||||||
|
|
||||||
- name: Enable and start docker-pull-images user timer
|
|
||||||
ansible.builtin.command: systemctl --user enable --now docker-pull-images.timer
|
|
||||||
become: true
|
|
||||||
become_user: "{{ gitea_runner_service_user }}"
|
|
||||||
environment:
|
|
||||||
XDG_RUNTIME_DIR: "/run/user/{{ gitea_runner_uid }}"
|
|
||||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ gitea_runner_uid | default(0) }}/bus"
|
|
||||||
changed_when: true
|
|
||||||
when:
|
|
||||||
- gitea_runner_systemd_available.stat.exists
|
|
||||||
- gitea_runner_docker_rootless_setup
|
|
||||||
- gitea_runner_pre_cache_schedule | length > 0
|
|
||||||
- gitea_runner_pre_cache_images | length > 0
|
|
||||||
|
|
||||||
- name: Disable and stop docker-pull-images timer (no images or schedule)
|
|
||||||
ansible.builtin.command: systemctl --user disable --now docker-pull-images.timer
|
|
||||||
become: true
|
|
||||||
become_user: "{{ gitea_runner_service_user }}"
|
|
||||||
environment:
|
|
||||||
XDG_RUNTIME_DIR: "/run/user/{{ gitea_runner_uid }}"
|
|
||||||
DBUS_SESSION_BUS_ADDRESS: "unix:path=/run/user/{{ gitea_runner_uid | default(0) }}/bus"
|
|
||||||
changed_when: true
|
|
||||||
failed_when: false
|
|
||||||
when:
|
|
||||||
- gitea_runner_systemd_available.stat.exists
|
|
||||||
- gitea_runner_docker_rootless_setup
|
|
||||||
- gitea_runner_pre_cache_schedule | length == 0 or gitea_runner_pre_cache_images | length == 0
|
|
||||||
@@ -8,15 +8,6 @@
|
|||||||
#
|
#
|
||||||
# Set gitea_runner_pre_pull_images to a list of image refs to pull, or
|
# Set gitea_runner_pre_pull_images to a list of image refs to pull, or
|
||||||
# empty list to skip pre-pulling.
|
# empty list to skip pre-pulling.
|
||||||
#
|
|
||||||
# IMPORTANT: Do NOT use this mechanism for:
|
|
||||||
# - CI runner container images (e.g. ci-full) — these are already
|
|
||||||
# cached by the runner setup task and pulling them here is redundant.
|
|
||||||
# - Images that molecule tests pull themselves — molecule prepare/converge
|
|
||||||
# steps handle their own image pulls; pre-pulling them here wastes time
|
|
||||||
# and disk space.
|
|
||||||
# This mechanism is intended only for images that are needed by the runner
|
|
||||||
# itself but not pulled by any molecule scenario or runner setup step.
|
|
||||||
|
|
||||||
- name: Pre-pull Docker images for CI runner
|
- name: Pre-pull Docker images for CI runner
|
||||||
ansible.builtin.command: "docker pull {{ item }}"
|
ansible.builtin.command: "docker pull {{ item }}"
|
||||||
|
|||||||
@@ -5,18 +5,16 @@ Description=Docker prune for Gitea runner resources
|
|||||||
Type=oneshot
|
Type=oneshot
|
||||||
Environment=DOCKER_HOST=unix:///run/user/{{ gitea_runner_uid }}/docker.sock
|
Environment=DOCKER_HOST=unix:///run/user/{{ gitea_runner_uid }}/docker.sock
|
||||||
Environment=XDG_RUNTIME_DIR=/run/user/{{ gitea_runner_uid }}
|
Environment=XDG_RUNTIME_DIR=/run/user/{{ gitea_runner_uid }}
|
||||||
# Force-remove stale *stopped* containers left behind by failed molecule tests.
|
# Force-remove stale containers (including running ones) left behind by failed
|
||||||
# Implements: REQ-1 (GRM-166) — only containers with status=exited are
|
# molecule tests. "docker container prune -f" only removes stopped containers,
|
||||||
# eligible. RunningFor measures creation time, so a stale molecule instance
|
# so running containers from crashed/interrupted CI jobs accumulate indefinitely,
|
||||||
# (e.g. ubuntu-2604) that a new run restarts still looks ">1h old"; removing
|
# consuming disk and memory. We stop+rm everything first, then prune the rest.
|
||||||
# running containers kills active converges with "No such container"
|
|
||||||
# (infra nightly run 5710). Running leftovers are instead reused or destroyed
|
|
||||||
# by the next molecule create/destroy cycle.
|
|
||||||
# Exclude CI job containers (name starts with GITEA-ACTIONS-TASK) — removing
|
# Exclude CI job containers (name starts with GITEA-ACTIONS-TASK) — removing
|
||||||
# them kills the active CI job and causes "RWLayer is unexpectedly nil" errors.
|
# them kills the active CI job and causes "RWLayer is unexpectedly nil" errors.
|
||||||
# Only remove containers older than 1 hour (grep for "hour/day/week/month/year
|
# Only remove containers older than 1 hour (grep for "hour/day/week/month/year
|
||||||
# ago" in RunningFor) to avoid removing containers a job just created.
|
# ago" in RunningFor) to avoid killing molecule test containers that CI jobs
|
||||||
ExecStart=/bin/sh -c 'docker ps -a --filter "status=exited" --format "{% raw %}{{.ID}} {{.Names}} {{.RunningFor}}{% endraw %}" 2>/dev/null | grep -v "GITEA-ACTIONS-TASK" | grep -E "(hour|day|week|month|year)s? ago" | awk "{print $1}" | xargs -r docker rm -f 2>/dev/null || true'
|
# are actively using.
|
||||||
|
ExecStart=/bin/sh -c 'docker ps -a --format "{% raw %}{{.ID}} {{.Names}} {{.RunningFor}}{% endraw %}" 2>/dev/null | grep -v "GITEA-ACTIONS-TASK" | grep -E "(hour|day|week|month|year)s? ago" | awk "{print $1}" | xargs -r docker rm -f 2>/dev/null || true'
|
||||||
ExecStart=/usr/bin/docker system prune -af --filter "until={{ gitea_runner_prune_until }}" --volumes
|
ExecStart=/usr/bin/docker system prune -af --filter "until={{ gitea_runner_prune_until }}" --volumes
|
||||||
# Prune networks older than the prune-until threshold to avoid removing
|
# Prune networks older than the prune-until threshold to avoid removing
|
||||||
# networks that molecule tests are actively creating (e.g. 'traefik' network
|
# networks that molecule tests are actively creating (e.g. 'traefik' network
|
||||||
|
|||||||
@@ -1,14 +0,0 @@
|
|||||||
[Unit]
|
|
||||||
Description=Pre-pull Docker images for CI runner cache
|
|
||||||
After=docker.service
|
|
||||||
Wants=docker.service
|
|
||||||
|
|
||||||
[Service]
|
|
||||||
Type=oneshot
|
|
||||||
Environment=DOCKER_HOST=unix:///run/user/{{ gitea_runner_uid }}/docker.sock
|
|
||||||
Environment=XDG_RUNTIME_DIR=/run/user/{{ gitea_runner_uid }}
|
|
||||||
# Pull each image quietly. docker pull exits 0 if image is already up-to-date,
|
|
||||||
# so this is idempotent. Errors are non-fatal (image may already be cached).
|
|
||||||
{% for image in gitea_runner_pre_cache_images %}
|
|
||||||
ExecStart=/usr/bin/docker pull -q {{ image }}
|
|
||||||
{% endfor %}
|
|
||||||
@@ -1,10 +0,0 @@
|
|||||||
[Unit]
|
|
||||||
Description=Periodic Docker image pre-cache for CI runner
|
|
||||||
|
|
||||||
[Timer]
|
|
||||||
OnCalendar={{ gitea_runner_pre_cache_schedule }}
|
|
||||||
Persistent=true
|
|
||||||
RandomizedDelaySec=300
|
|
||||||
|
|
||||||
[Install]
|
|
||||||
WantedBy=timers.target
|
|
||||||
@@ -9,13 +9,3 @@ runner:
|
|||||||
container:
|
container:
|
||||||
label: "{{ gitea_runner_container_label }}"
|
label: "{{ gitea_runner_container_label }}"
|
||||||
docker_host: "unix:///run/user/{{ gitea_runner_uid }}/docker.sock"
|
docker_host: "unix:///run/user/{{ gitea_runner_uid }}/docker.sock"
|
||||||
force_pull: {{ gitea_runner_force_pull | lower }}
|
|
||||||
{% if gitea_runner_container_options | length > 0 %}
|
|
||||||
options: "{{ gitea_runner_container_options }}"
|
|
||||||
{% endif %}
|
|
||||||
{% if gitea_runner_valid_volumes | length > 0 %}
|
|
||||||
valid_volumes:
|
|
||||||
{% for volume in gitea_runner_valid_volumes %}
|
|
||||||
- "{{ volume }}"
|
|
||||||
{% endfor %}
|
|
||||||
{% endif %}
|
|
||||||
|
|||||||
@@ -30,50 +30,6 @@ if [[ -n "$stuck_containers" ]]; then
|
|||||||
echo "$stuck_containers" | xargs -r docker rm -f 2>/dev/null || true
|
echo "$stuck_containers" | xargs -r docker rm -f 2>/dev/null || true
|
||||||
fi
|
fi
|
||||||
|
|
||||||
# 1c. Detect stalled CI job containers — Implements: REQ-1..REQ-4 (GRM-168)
|
|
||||||
# act_runner exec/archive calls into long-running job containers have
|
|
||||||
# repeatedly timed out ~20min into jobs while the daemon stayed up.
|
|
||||||
# Probe exec responsiveness on aged job containers and, on timeout,
|
|
||||||
# write one diagnostics bundle per container for post-mortem analysis.
|
|
||||||
STALL_MINUTES={{ gitea_runner_stall_minutes }}
|
|
||||||
DIAG_DIR="{{ gitea_runner_config_dir }}"
|
|
||||||
now_epoch=$(date +%s)
|
|
||||||
# Implements: REQ-1 — guard the enumeration: a slow/dead daemon must not
|
|
||||||
# abort the healthcheck under pipefail; an empty list just skips probing.
|
|
||||||
# Implements: REQ-2 — pipe-separate fields: CreatedAt contains spaces, so
|
|
||||||
# whitespace-splitting `read` only captured the date and broke the age gate.
|
|
||||||
{ timeout 15 docker ps --filter "name=GITEA-ACTIONS-TASK" \
|
|
||||||
--format '{% raw %}{{.ID}}|{{.Names}}|{{.CreatedAt}}{% endraw %}' 2>/dev/null || true; } \
|
|
||||||
| while IFS='|' read -r cid cname ccreated _rest; do
|
|
||||||
# GNU date rejects the redundant " +0000 UTC" suffix — drop it.
|
|
||||||
created_epoch=$(date -d "${ccreated% UTC}" +%s 2>/dev/null || echo 0)
|
|
||||||
age_min=$(( (now_epoch - created_epoch) / 60 ))
|
|
||||||
[[ "$age_min" -lt "$STALL_MINUTES" ]] && continue
|
|
||||||
marker="$DIAG_DIR/.stall-diag-$cid"
|
|
||||||
[[ -f "$marker" ]] && continue
|
|
||||||
if ! timeout 10 docker exec "$cid" true 2>/dev/null; then
|
|
||||||
diag="$DIAG_DIR/stall-diag-$cname-$(date +%Y%m%dT%H%M%S).log"
|
|
||||||
{
|
|
||||||
echo "=== stall diagnostics for $cname ($cid), age ${age_min}m ==="
|
|
||||||
echo "--- exec probe: TIMEOUT (>10s) ---"
|
|
||||||
echo "--- docker inspect ---"
|
|
||||||
# Implements: REQ-3 — full inspect, but redact the Env block:
|
|
||||||
# job containers carry CI tokens in env vars; the bundle must
|
|
||||||
# not become a secret-material artifact.
|
|
||||||
timeout 15 docker inspect "$cid" 2>/dev/null \
|
|
||||||
| sed -E 's/("[^"]*(TOKEN|PASSWORD|SECRET|KEY)[^=]*=)[^",]*/\1<redacted>/Ig'
|
|
||||||
echo "--- docker top ---"
|
|
||||||
timeout 15 docker top "$cid" 2>/dev/null
|
|
||||||
echo "--- docker stats --no-stream ---"
|
|
||||||
timeout 15 docker stats --no-stream "$cid" 2>/dev/null
|
|
||||||
echo "--- docker events --since 30m ---"
|
|
||||||
timeout 15 docker events --since 30m --until 0s 2>/dev/null | tail -50
|
|
||||||
} > "$diag" 2>&1 || true
|
|
||||||
touch "$marker"
|
|
||||||
echo "WARN: job container $cname unresponsive to exec (${age_min}m old) — diagnostics at $diag"
|
|
||||||
fi
|
|
||||||
done
|
|
||||||
|
|
||||||
# 2. Check gitea-runner service is active
|
# 2. Check gitea-runner service is active
|
||||||
runner_state=$(systemctl --user is-active gitea-runner.service 2>/dev/null || true)
|
runner_state=$(systemctl --user is-active gitea-runner.service 2>/dev/null || true)
|
||||||
if [[ "$runner_state" != "active" ]]; then
|
if [[ "$runner_state" != "active" ]]; then
|
||||||
@@ -271,12 +227,9 @@ if [[ "$disk_pct" -ge {{ gitea_runner_healthcheck_disk_critical }} ]]; then
|
|||||||
echo "CRITICAL: Disk usage at ${disk_pct}% (>= {{ gitea_runner_healthcheck_disk_critical }}%), full prune"
|
echo "CRITICAL: Disk usage at ${disk_pct}% (>= {{ gitea_runner_healthcheck_disk_critical }}%), full prune"
|
||||||
# Critical level: remove ALL stopped containers (no age filter) and ALL
|
# Critical level: remove ALL stopped containers (no age filter) and ALL
|
||||||
# unused images/volumes. The until=1h gentle prune is insufficient here.
|
# unused images/volumes. The until=1h gentle prune is insufficient here.
|
||||||
# Implements: REQ-1 (GRM-167) — only exited/dead containers are removed.
|
# Stop+rm stale non-CI containers regardless of age (failed molecule tests
|
||||||
# Running molecule instances are never killed: RunningFor counts creation
|
# from the last 59 minutes also consume disk).
|
||||||
# time, so an adopted stale instance looks old; and a running container's
|
docker ps -a --format '{% raw %}{{.ID}} {{.Names}}{% endraw %}' 2>/dev/null \
|
||||||
# writable layer is tiny — images/volumes are what actually fills the disk.
|
|
||||||
docker ps -a --filter "status=exited" --filter "status=dead" \
|
|
||||||
--format '{% raw %}{{.ID}} {{.Names}}{% endraw %}' 2>/dev/null \
|
|
||||||
| grep -v 'GITEA-ACTIONS-TASK' \
|
| grep -v 'GITEA-ACTIONS-TASK' \
|
||||||
| awk '{print $1}' \
|
| awk '{print $1}' \
|
||||||
| xargs -r docker rm -f 2>/dev/null || true
|
| xargs -r docker rm -f 2>/dev/null || true
|
||||||
@@ -287,16 +240,13 @@ if [[ "$disk_pct" -ge {{ gitea_runner_healthcheck_disk_critical }} ]]; then
|
|||||||
echo "INFO: Disk usage after full prune: ${disk_pct}%"
|
echo "INFO: Disk usage after full prune: ${disk_pct}%"
|
||||||
elif [[ "$disk_pct" -ge {{ gitea_runner_healthcheck_disk_threshold }} ]]; then
|
elif [[ "$disk_pct" -ge {{ gitea_runner_healthcheck_disk_threshold }} ]]; then
|
||||||
echo "WARN: Disk usage at ${disk_pct}%, pruning runner resources (until=1h)"
|
echo "WARN: Disk usage at ${disk_pct}%, pruning runner resources (until=1h)"
|
||||||
# Force-remove stale stopped containers older than 1 hour.
|
# Force-remove stale containers (including running ones from failed molecule tests)
|
||||||
# Implements: REQ-1 (GRM-167) — only exited/dead containers are removed.
|
# that are older than 1 hour. "docker container prune -f" only removes stopped
|
||||||
# A running molecule instance must never be janitor-killed: RunningFor
|
# containers, so running containers from crashed CI jobs accumulate and consume
|
||||||
# measures creation time, so a stale instance restarted by an active run
|
# disk/memory. Exclude CI job containers (name starts with GITEA-ACTIONS-TASK).
|
||||||
# looks ">1h old" and would die mid-converge ("No such container",
|
# Only remove containers older than 1 hour to avoid killing molecule test
|
||||||
# infra nightly run 5710). Running leftovers are reused or destroyed by
|
# containers that CI jobs are actively using.
|
||||||
# the next molecule create/destroy cycle.
|
docker ps -a --format '{% raw %}{{.ID}} {{.Names}} {{.RunningFor}}{% endraw %}' 2>/dev/null \
|
||||||
# Exclude CI job containers (name starts with GITEA-ACTIONS-TASK).
|
|
||||||
docker ps -a --filter "status=exited" --filter "status=dead" \
|
|
||||||
--format '{% raw %}{{.ID}} {{.Names}} {{.RunningFor}}{% endraw %}' 2>/dev/null \
|
|
||||||
| grep -v 'GITEA-ACTIONS-TASK' \
|
| grep -v 'GITEA-ACTIONS-TASK' \
|
||||||
| grep -E '(hour|day|week|month|year)s? ago' \
|
| grep -E '(hour|day|week|month|year)s? ago' \
|
||||||
| awk '{print $1}' \
|
| awk '{print $1}' \
|
||||||
|
|||||||
+6
-6
@@ -8,12 +8,12 @@ Each runner runs in an isolated **rootless Docker** environment under a dedicate
|
|||||||
|
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/src/branch/master/LICENSE)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/wiki)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/actions)
|
||||||
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases)
|
[](https://git.oblachno.oblachno.fyi/oblachno-oss/grm/releases)
|
||||||
[](https://www.python.org/downloads/)
|
[](https://www.python.org/downloads/)
|
||||||
|
|
||||||
## Overview
|
## Overview
|
||||||
|
|
||||||
|
|||||||
@@ -1,47 +0,0 @@
|
|||||||
# GRM-162: Add pre-cache timer, force_pull, and Docker socket options to runner config
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
CI containers were not using the host's rootless Docker daemon, leading to
|
|
||||||
"no space left on device" errors. The runner config template was missing
|
|
||||||
`force_pull`, `options` (host Docker socket mount), and `valid_volumes`
|
|
||||||
fields. Additionally, no pre-cache timer existed to prevent thundering-herd
|
|
||||||
registry timeouts when all runners pull images simultaneously.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
REQ-1: Add `force_pull: false` to runner config template (explicit default
|
|
||||||
so the runner reuses locally cached images instead of pulling on every job)
|
|
||||||
REQ-2: Add `options` field to mount host rootless Docker socket as
|
|
||||||
`/run/host-docker.sock` so `start_docker.py` inside CI containers can detect
|
|
||||||
and use the host daemon (full disk, no nested DinD)
|
|
||||||
REQ-3: Add `valid_volumes` list for the socket mount targets (validated by
|
|
||||||
the runner against `container.options` and job-level volumes)
|
|
||||||
REQ-4: Add `pre_cache.yml` task with a systemd user timer that pre-pulls CI
|
|
||||||
images every 6 hours (configurable via `gitea_runner_pre_cache_schedule`)
|
|
||||||
REQ-5: Add `docker-pull-images.service.j2` and `docker-pull-images.timer.j2`
|
|
||||||
templates for the pre-cache timer
|
|
||||||
REQ-6: Timer is disabled when `gitea_runner_pre_cache_schedule` is empty or
|
|
||||||
`gitea_runner_pre_cache_images` is empty (graceful degradation)
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
- `make lint-ci` passes (ansible-lint on new task/template files)
|
|
||||||
- `make molecule` converges successfully with the new pre-cache tasks
|
|
||||||
- Verify the runner config template renders correctly with and without
|
|
||||||
container options/valid_volumes
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
- Merge to master → post-merge auto-publishes package
|
|
||||||
- Infra dependency PR auto-created to bump pinned grm version
|
|
||||||
- Runners pick up the new config on next `make setup` or ansible apply
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
- Revert the merge commit
|
|
||||||
- Set `gitea_runner_pre_cache_schedule: ""` to disable the timer without
|
|
||||||
reverting
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
- [x] REQ-1: `force_pull: false` in runner config template
|
|
||||||
- [x] REQ-2: `options` field mounts host Docker socket as `/run/host-docker.sock`
|
|
||||||
- [x] REQ-3: `valid_volumes` list includes both socket mount targets
|
|
||||||
- [x] REQ-4: `pre_cache.yml` task creates and manages systemd user timer
|
|
||||||
- [x] REQ-5: Service and timer templates created
|
|
||||||
- [x] REQ-6: Timer disabled gracefully when schedule or images empty
|
|
||||||
@@ -1,43 +0,0 @@
|
|||||||
# GRM-166: Fix docker-prune killing running molecule containers
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
`docker-prune.service` force-removes containers older than 1h by
|
|
||||||
**creation** time (`RunningFor`). Molecule instance containers
|
|
||||||
(`ubuntu-2604` etc.) run on the runner's shared rootless daemon and are
|
|
||||||
not name-excluded (only `GITEA-ACTIONS-TASK` jobs are). A stale molecule
|
|
||||||
container left by a crashed run is *restarted/reused* by the next run's
|
|
||||||
create phase — it then shows `RunningFor` >1h and gets `docker rm -f`'d
|
|
||||||
mid-converge: `No such container: ubuntu-2604`. Observed on infra nightly
|
|
||||||
run 5710 across multiple runners (incl. a 32 GB host), causing
|
|
||||||
shard failures.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
REQ-1: Change the container cleanup `ExecStart` in
|
|
||||||
`ansible/roles/gitea_runner/templates/docker-prune.service.j2` to only
|
|
||||||
target containers with `status=exited` (add
|
|
||||||
`--filter "status=exited"`). Running containers — including stale
|
|
||||||
molecule instances adopted by an active run — are never force-removed.
|
|
||||||
REQ-2: Update the template comment to document the race and why only
|
|
||||||
stopped containers are removed.
|
|
||||||
REQ-3: Extend the `template-content` molecule verify assertions to
|
|
||||||
require the `status=exited` filter in the rendered unit.
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
- `make molecule` (template-content scenario) verifies the rendered unit.
|
|
||||||
- `make pytest-cov` and `make lint-all` pass.
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
- Merge to master → package publishes; runner hosts pick up the role on
|
|
||||||
their next grm install/upgrade cycle (or manual re-run of the role on
|
|
||||||
affected runners).
|
|
||||||
- Until then, nightly molecule reruns remain exposed to the race —
|
|
||||||
mitigated by the fact that swept daemons have no stale containers left.
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
- Revert the merge commit. Running stale leftovers would then again be
|
|
||||||
force-removed (the pre-existing risky behaviour).
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
- [x] REQ-1: prune rm step filtered to `status=exited`
|
|
||||||
- [x] REQ-2: comment documents the creation-time vs running race
|
|
||||||
- [x] REQ-3: template-content verify asserts the exited filter
|
|
||||||
@@ -1,21 +0,0 @@
|
|||||||
# GRM-167: Bump devx to v0.51.0
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
devx v0.51.0 released with role defaults path support for create_dependency_pr. grm pins v0.50.2.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
Bump devx pin in pyproject.toml.
|
|
||||||
|
|
||||||
REQ-1: Bump devx from v0.50.2 to v0.51.0 in pyproject.toml
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
- Verify CI passes with new devx version
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
- Merge to master, auto-release
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
- Revert the merge commit
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
- [x] REQ-1: devx pinned to v0.51.0 in pyproject.toml
|
|
||||||
@@ -1,51 +0,0 @@
|
|||||||
# GRM-167: Fix runner healthcheck force-removing active molecule containers
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
`runner-healthcheck.sh` runs every 2 minutes and force-removes molecule
|
|
||||||
instance containers on two disk paths:
|
|
||||||
|
|
||||||
- **Warn path (disk >= 70%)**: `docker rm -f` any non-`GITEA-ACTIONS-TASK`
|
|
||||||
container with `RunningFor` >= 1h — same creation-age race fixed in
|
|
||||||
GRM-166 for the prune service. A stale molecule instance restarted by a
|
|
||||||
new run still reads >1h old and is killed mid-converge.
|
|
||||||
- **Critical path (disk >= 75%)**: `docker rm -f` ALL
|
|
||||||
non-`GITEA-ACTIONS-TASK` containers with no age or status filter —
|
|
||||||
running molecule instances die instantly. Under parallel DinD load the
|
|
||||||
runner disk crosses 75% routinely (infra nightly run 5710: `No such
|
|
||||||
container: ubuntu-2604` across three different runners).
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
REQ-1: In `ansible/roles/gitea_runner/templates/runner-healthcheck.sh.j2`,
|
|
||||||
restrict both disk-pressure removal paths to containers that are not
|
|
||||||
running: add `--filter "status=exited" --filter "status=created"` is
|
|
||||||
NOT sufficient for the critical path since a just-created molecule
|
|
||||||
instance is in `created` state — use `status=exited` and
|
|
||||||
`status=dead` only. Running containers are never janitor-killed; the
|
|
||||||
image/volume/system prunes still reclaim the actual disk.
|
|
||||||
REQ-2: Keep `GITEA-ACTIONS-TASK` exclusion, the 1h age gate on the warn
|
|
||||||
path, and all image/volume/network/builder prunes unchanged.
|
|
||||||
REQ-3: Update comments documenting the running-container race.
|
|
||||||
REQ-4: Extend molecule `template-content` verify assertions for the
|
|
||||||
rendered healthcheck script (`status=exited` present in both paths).
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
- `make molecule` template-content + default scenarios pass.
|
|
||||||
- `make pytest-cov`, `make lint-all` pass.
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
- Merge to master → package publishes. Runner hosts apply the role on
|
|
||||||
their next grm install/upgrade run; affected runners may need a manual
|
|
||||||
role re-run for immediate relief.
|
|
||||||
- Vikunja counter reused GRM-167; prior spec preserved in
|
|
||||||
[GRM-167-devx-bump-historical](GRM-167-devx-bump-historical.md).
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
- Revert the merge commit — restores the aggressive cleanup that kills
|
|
||||||
active CI jobs.
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
- [x] REQ-1: both disk paths only remove `status=exited`/`status=dead`
|
|
||||||
containers
|
|
||||||
- [x] REQ-2: exclusions, age gate, and non-container prunes unchanged
|
|
||||||
- [x] REQ-3: comments updated
|
|
||||||
- [x] REQ-4: template-content verify covers the filters
|
|
||||||
@@ -1,84 +0,0 @@
|
|||||||
# GRM-168: Capture dockerd diagnostics when a CI job container stalls
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
|
|
||||||
Recurring CI failures (5+ times on 2026-09-17/18, infra runs 5913, 5925,
|
|
||||||
5943, 5955 notify-sso-bridge): ~20 min into a long-running job,
|
|
||||||
act_runner's API calls into the job container (`docker exec`, archive
|
|
||||||
fetch of `/var/run/act/workflow/*.txt`) time out with
|
|
||||||
`docker daemon ping during version negotiation failed /
|
|
||||||
context deadline exceeded` — killing the job.
|
|
||||||
|
|
||||||
Established facts:
|
|
||||||
|
|
||||||
- Host rootless dockerd never restarted (all daemons up since Sep 14);
|
|
||||||
the healthcheck's 10 s `docker info` never timed out — the daemon API
|
|
||||||
stayed responsive at daemon level.
|
|
||||||
- No OOM, disk, inode, or load pressure on the host.
|
|
||||||
- The wedge is therefore per-container (shim/exec path), most consistent
|
|
||||||
with attach-stdio backpressure or a containerd-shim event stall — but
|
|
||||||
cannot be confirmed post-mortem because job containers and their
|
|
||||||
dockerd goroutine state are gone by the time anyone looks.
|
|
||||||
|
|
||||||
A `SIGUSR1` dockerd dump is not useful here: it lands in the user
|
|
||||||
journal, which runner users cannot read (2026-08-08 journal-permission
|
|
||||||
incident documented in this file's header comments).
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
|
|
||||||
Extend `runner-healthcheck.sh.j2` with a stall-detection section that
|
|
||||||
runs after the daemon liveness check. On every healthcheck tick (2 min):
|
|
||||||
|
|
||||||
REQ-1: For each running `GITEA-ACTIONS-TASK-*` container older than
|
|
||||||
`gitea_runner_stall_minutes` (default 15), probe exec responsiveness
|
|
||||||
with `timeout 10 docker exec <id> true`.
|
|
||||||
|
|
||||||
REQ-2: If the probe times out, write a diagnostics bundle to
|
|
||||||
`{{ gitea_runner_config_dir }}/stall-diag-<container>-<timestamp>.log`
|
|
||||||
containing: probe result, `docker inspect` output (State, OOMKilled,
|
|
||||||
Pid, finished/started times), `docker top` output, `docker stats
|
|
||||||
--no-stream` for the container, and `docker events --since 30m` output.
|
|
||||||
Each line prefixed with the container name for grepability.
|
|
||||||
|
|
||||||
REQ-3: Cooldown per container — write at most one diagnostics bundle
|
|
||||||
per container id (marker file under the same dir), so a 2-minute
|
|
||||||
healthcheck does not spam dumps on a persistent stall.
|
|
||||||
|
|
||||||
REQ-4: Do not kill or restart anything — diagnostics only. The job may
|
|
||||||
recover on its own; if it does not, the captured evidence isolates
|
|
||||||
shim-vs-daemon and stream-vs-exec for the follow-up fix.
|
|
||||||
|
|
||||||
## Files Affected
|
|
||||||
|
|
||||||
- `ansible/roles/gitea_runner/templates/runner-healthcheck.sh.j2` (extend)
|
|
||||||
- `ansible/roles/gitea_runner/defaults/main.yml` (add `gitea_runner_stall_minutes`)
|
|
||||||
- `docs/specs/GRM-168.md` (new)
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
|
|
||||||
- `make lint-all` (ansible-lint + shellcheck-adjacent linters) passes.
|
|
||||||
- Molecule fast-converge on the gitea_runner role scenario that deploys
|
|
||||||
the healthcheck template (template renders without error).
|
|
||||||
- Manual trace: the new section only touches containers matching
|
|
||||||
`GITEA-ACTIONS-TASK-*` older than the threshold; a stalled exec probe
|
|
||||||
writes exactly one bundle per container.
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
|
|
||||||
Merge via auto-merge → GRM release → infra picks up the new version via
|
|
||||||
the automated dependency PR. Runner hosts get the updated healthcheck on
|
|
||||||
the next `gitea_runner` role apply (nightly or manual run).
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
|
|
||||||
Revert the template change — the healthcheck returns to the previous
|
|
||||||
probe set. The diagnostics path is additive; removing it risks nothing.
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
|
|
||||||
- [x] Stalled job containers probed via `timeout docker exec`.
|
|
||||||
- [x] One diagnostics bundle per stalled container, written to the
|
|
||||||
runner config dir (readable without journal access).
|
|
||||||
- [x] Per-container cooldown prevents dump spam.
|
|
||||||
- [x] Nothing is killed/restarted — diagnostics only.
|
|
||||||
- [x] `make lint-all` passes.
|
|
||||||
@@ -1,48 +0,0 @@
|
|||||||
# GRM-169: Fix auto-merge timeout — molecule wait exceeds 10-min job cap
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
|
|
||||||
The `auto-merge` job in `ci.yml` polls molecule-tests status with
|
|
||||||
`MAX_WAIT=600` (10 minutes) inside a job capped at
|
|
||||||
`timeout-minutes: 10`. The molecule suite takes ~25 minutes under the
|
|
||||||
4-runner distribution. Result: every `pull_request` synchronize run of
|
|
||||||
auto-merge exhausts MAX_WAIT, prints `Timed out waiting for molecule
|
|
||||||
tests`, and fails — observed on PR #275 (run 5988) where all molecule
|
|
||||||
jobs were green but auto-merge died before they finished. The merge
|
|
||||||
only completed via a manual rerun-failed-jobs call after molecule was
|
|
||||||
already green.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
|
|
||||||
REQ-1: Raise the molecule wait budget in `.gitea/workflows/ci.yml` so it
|
|
||||||
exceeds the observed suite duration: `MAX_WAIT=2400` (40 minutes — ~1.6x
|
|
||||||
the observed 25-minute suite) and the job `timeout-minutes` to `50`
|
|
||||||
(wait budget plus setup/post overhead).
|
|
||||||
|
|
||||||
REQ-2: No other behavior changes — the wait loop, success/failure/skipped
|
|
||||||
classification, and merge semantics stay identical. The job still runs
|
|
||||||
on every pull_request event; it simply no longer aborts early.
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
|
|
||||||
- `make workflow-lint` (actionlint) passes on the edited file.
|
|
||||||
- `make workflow-dryrun` where available.
|
|
||||||
- Next PR's auto-merge run waits past the 10-minute mark and merges
|
|
||||||
after molecule turns green (verified on a subsequent PR).
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
|
|
||||||
Merge via auto-merge — ironically exercised by this very PR's auto-merge
|
|
||||||
run: it must wait for this PR's own molecule jobs, demonstrating the fix
|
|
||||||
in production immediately.
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
|
|
||||||
Revert the two changed lines. Risk of keeping the fix: none — a longer
|
|
||||||
wait can only extend a job that was previously guaranteed to fail.
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
|
|
||||||
- [x] `MAX_WAIT` raised to 2400 in the auto-merge wait loop.
|
|
||||||
- [x] `timeout-minutes` raised to 50 on the auto-merge job.
|
|
||||||
- [x] `make workflow-lint` passes.
|
|
||||||
@@ -1,73 +0,0 @@
|
|||||||
# GRM-170: Fix stall-detection robustness bugs in runner healthcheck
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
|
|
||||||
Post-merge review of the GRM-168 stall-detection block in
|
|
||||||
`runner-healthcheck.sh.j2` found three defects:
|
|
||||||
|
|
||||||
1. **Missing `|| true` on the container enumeration.** `timeout 15
|
|
||||||
docker ps … | while …` runs under `set -euo pipefail`. If the daemon
|
|
||||||
is unresponsive — precisely the condition the section exists to
|
|
||||||
diagnose — `docker ps` exits nonzero, pipefail propagates it, and the
|
|
||||||
healthcheck dies mid-run before reaching the runner-service check.
|
|
||||||
Every other docker call in the script is guarded; this one is not.
|
|
||||||
|
|
||||||
2. **CreatedAt split bug.** `docker ps --format '{{.ID}} {{.Names}}
|
|
||||||
{{.CreatedAt}}'` emits a timestamp containing spaces
|
|
||||||
(`2026-09-18 10:30:00 +0000 UTC`), but `read -r cid cname ccreated
|
|
||||||
_rest` only captures `2026-09-18` — the date part. `date -d` then
|
|
||||||
computes age from midnight: containers created today always appear
|
|
||||||
≥N hours old, so the 15-minute gate effectively never filters.
|
|
||||||
|
|
||||||
3. **`head -200` truncates `docker inspect`.** Inspect output is ~300+
|
|
||||||
lines and the `State` block (OOMKilled, Pid, times) the spec requires
|
|
||||||
can be cut off.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
|
|
||||||
REQ-1: Wrap the enumeration so a failed `docker ps` yields empty input
|
|
||||||
instead of aborting the script: `{ timeout 15 docker ps … || true; } |
|
|
||||||
while …`.
|
|
||||||
|
|
||||||
REQ-2: Emit fields separated by `|` (`{{.ID}}|{{.Names}}|{{.CreatedAt}}`)
|
|
||||||
and parse with `IFS='|' read -r cid cname ccreated _rest` so the full
|
|
||||||
timestamp reaches `date -d`; also strip the redundant ` UTC` suffix
|
|
||||||
because GNU date rejects `+0000 UTC` together. The age gate then
|
|
||||||
compares real minutes.
|
|
||||||
|
|
||||||
REQ-3: Remove the `head -200` truncation on `docker inspect` output so
|
|
||||||
the full State block is captured — but pipe through a `sed` filter that
|
|
||||||
redacts the value of any env entry whose name contains TOKEN, PASSWORD,
|
|
||||||
SECRET, or KEY. Job containers carry CI tokens in their Env block; the
|
|
||||||
diagnostics bundle must not become a secret-material artifact
|
|
||||||
(OBL-INFRA-548 S02).
|
|
||||||
|
|
||||||
REQ-4: Diagnostics-only constraint unchanged — no kills, no restarts.
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
|
|
||||||
- Render the template and run `bash -n` on the output.
|
|
||||||
- Shell-simulate: feed a fake `docker ps` line with spaced CreatedAt and
|
|
||||||
verify `date -d` computes minutes correctly (manual check).
|
|
||||||
- `make lint-all` (ansible-lint, actionlint, ruff) passes.
|
|
||||||
- Molecule gitea_runner scenario converges with the template change.
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
|
|
||||||
Merge via auto-merge → release (fix: commit bumps patch) → infra
|
|
||||||
dependency-bump PR picks up the new role version → runner role applied
|
|
||||||
on next infra run. This PR also carries the merged-but-unreleased
|
|
||||||
GRM-168 healthcheck into the release.
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
|
|
||||||
Revert the three-line change set; the section degrades to the GRM-168
|
|
||||||
behavior (still diagnostics-only, just less robust).
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
|
|
||||||
- [x] `docker ps` enumeration guarded against nonzero exit.
|
|
||||||
- [x] Full CreatedAt timestamp parsed via `|` separator.
|
|
||||||
- [x] `docker inspect` captured without truncation.
|
|
||||||
- [x] Rendered script passes `bash -n`.
|
|
||||||
- [x] `make lint-all` passes.
|
|
||||||
@@ -1,28 +0,0 @@
|
|||||||
# GRM-171: Use kireto token for auto-merge approval review
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
The auto-merge workflow posts approval reviews with
|
|
||||||
`REVIEWER_GITEA_API_TOKEN` (emil), but emil is also the PR creator.
|
|
||||||
Gitea ignores self-approvals, so the merge fails with HTTP 405
|
|
||||||
`Does not have enough approvals`.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
REQ-1: Change the approval review step in `.gitea/workflows/ci.yml` to use
|
|
||||||
`DEVELOPER_GITEA_API_TOKEN` (kireto) instead of
|
|
||||||
`REVIEWER_GITEA_API_TOKEN` (emil), since kireto is a different user
|
|
||||||
than the PR creator.
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
- `make lint-all` passes (workflow-lint validates the YAML)
|
|
||||||
- Next auto-merge PR succeeds (approval posted by kireto, merge completes)
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
- Merge to master
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
- Revert the merge commit
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
- [x] REQ-1: Change the approval review step in `.gitea/workflows/ci.yml`
|
|
||||||
to use `DEVELOPER_GITEA_API_TOKEN` (kireto) instead of
|
|
||||||
`REVIEWER_GITEA_API_TOKEN` (emil)
|
|
||||||
@@ -1,51 +0,0 @@
|
|||||||
# GRM-171: Add runner-ops, molecule-testing, vikunja-tasks skills, fix create-task docs
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
|
|
||||||
The OBL-INFRA-548 programme audit found grm lacks skills for runner
|
|
||||||
fleet operations (needed for S08: leases, admission, watermarks),
|
|
||||||
molecule scenario authoring, and Vikunja task lifecycle.
|
|
||||||
`devx-workflow` and `AGENTS.md` document `make create-task -- --title`,
|
|
||||||
which fails because `devx-create-task` forwards no arguments.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
|
|
||||||
REQ-1: Add `runner-ops` skill: RunnerManager/AnsibleExecutor model,
|
|
||||||
stale-runner cleanup, image pruning, molecule container lifecycle,
|
|
||||||
safe-debugging rules.
|
|
||||||
REQ-2: Add `molecule-testing` skill: gitea_runner scenario layout,
|
|
||||||
platform overrides, isolation flags, debugging.
|
|
||||||
REQ-3: Add `vikunja-tasks` skill: create via module call, query,
|
|
||||||
close, spec-collision convention.
|
|
||||||
REQ-4: Fix broken `make create-task -- --title` documentation in
|
|
||||||
`devx-workflow` skill and `AGENTS.md`.
|
|
||||||
REQ-5: Add skill validation tests (`tests/unit/test_skills_validation.py`)
|
|
||||||
+ fix stale make-target refs and missing sections in existing skills.
|
|
||||||
|
|
||||||
Preserve the colliding spec as
|
|
||||||
[GRM-171-kireto-token-historical](GRM-171-kireto-token-historical.md).
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
|
|
||||||
- `pytest tests/unit/test_skills_validation.py` passes (12 tests).
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
|
|
||||||
Documentation/skills only — auto-merge to master; no runtime deploy.
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
|
|
||||||
Revert the squash-merge commit; skills are inert documentation.
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
|
|
||||||
- [x] REQ-1: `runner-ops` skill exists.
|
|
||||||
- [x] REQ-2: `molecule-testing` skill exists.
|
|
||||||
- [x] REQ-3: `vikunja-tasks` skill exists.
|
|
||||||
- [x] REQ-4: create-task docs corrected.
|
|
||||||
- [x] REQ-5: Skill validation tests added and passing.
|
|
||||||
|
|
||||||
## Out of Scope
|
|
||||||
|
|
||||||
- Runner lease/admission implementation (S08 scope).
|
|
||||||
- Fixing `devx-create-task` argument forwarding (devx repo, S11).
|
|
||||||
@@ -1,38 +0,0 @@
|
|||||||
# GRM-172: Audit and document pre-pull image usage guidelines
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
The grm repo contains a runner-level `pre_pull_images.yml` task file that
|
|
||||||
pre-pulls Docker images to avoid repeated pulls on every CI run. However,
|
|
||||||
there was no audit confirming that molecule `prepare.yml` files are not
|
|
||||||
also redundantly pre-pulling images that the runner setup already caches.
|
|
||||||
Wasteful pre-pulling wastes CI time and disk space.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
Audit all molecule `prepare.yml` files in the grm repo for pre-pull tasks.
|
|
||||||
The audit found NO molecule prepare.yml files contain pre-pull tasks, so no
|
|
||||||
code removal is needed. Document the audit findings in a spec and add a
|
|
||||||
comment to the runner-level `pre_pull_images.yml` task file clarifying that
|
|
||||||
it should not be used for images that molecule tests pull themselves (to
|
|
||||||
avoid redundant pulls).
|
|
||||||
|
|
||||||
REQ-1: Audit all molecule prepare.yml files for pre-pull tasks and confirm none exist
|
|
||||||
REQ-2: Add documentation comment to pre_pull_images.yml stating it should not be used for CI runner container images (already cached by runner setup) or images molecule tests pull themselves
|
|
||||||
REQ-3: Confirm gitea_runner_pre_pull_images default remains empty ([]) which is correct
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
- Grep all molecule prepare.yml files for pre-pull patterns confirms zero matches
|
|
||||||
- Verify pre_pull_images.yml comment is present and accurate
|
|
||||||
- Verify gitea_runner_pre_pull_images default is [] in defaults/main.yml
|
|
||||||
- Run make lint-ci to confirm no lint regressions
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
- Merge to master via auto-merge workflow
|
|
||||||
- No runtime changes; documentation-only
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
- Revert the merge commit; comments are removed, no functional impact
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
- [x] REQ-1: No molecule prepare.yml files in the grm repo contain pre-pull tasks (audit confirmed via grep)
|
|
||||||
- [x] REQ-2: pre_pull_images.yml contains a comment documenting it should not be used for CI runner container images or images molecule tests pull themselves
|
|
||||||
- [x] REQ-3: gitea_runner_pre_pull_images default remains empty ([]) in defaults/main.yml
|
|
||||||
@@ -1,46 +0,0 @@
|
|||||||
# GRM-173: Add dependency-graph, deployment-coordination, and skill-creation skills
|
|
||||||
|
|
||||||
## Problem
|
|
||||||
Agents working across the oblachno ecosystem lack shared, written context
|
|
||||||
for three recurring struggles: (1) knowing which repo produces what and
|
|
||||||
the correct order for cross-repo changes, (2) coordinating grm releases
|
|
||||||
with the downstream infra dependency PR, and (3) creating and validating
|
|
||||||
new Devin skills consistently. Without these skills, agents repeatedly
|
|
||||||
make mistakes such as deploying infra before the grm dependency PR is
|
|
||||||
merged, or writing skills that fail the validator.
|
|
||||||
|
|
||||||
## Approach
|
|
||||||
Add three skill files under `.devin/skills/`. Two are shared skills
|
|
||||||
(`dependency-graph`, `skill-creation`) that must be identical across
|
|
||||||
repos; one is grm-specific (`deployment-coordination`). All three
|
|
||||||
follow the standard skill structure (H1 title, When to Invoke,
|
|
||||||
Prerequisites, core content) and reference real make targets, file
|
|
||||||
paths, and API endpoints.
|
|
||||||
|
|
||||||
REQ-1: Add `.devin/skills/dependency-graph/SKILL.md` — shared skill mapping the oblachno ecosystem (repos, produces/consumers, dependency chain, correct change order, state verification)
|
|
||||||
REQ-2: Add `.devin/skills/deployment-coordination/SKILL.md` — grm-specific skill covering release flow, downstream consumer, coordinating a grm change, and common mistakes
|
|
||||||
REQ-3: Add `.devin/skills/skill-creation/SKILL.md` — shared skill for creating, validating, and maintaining skills (structure, quality standards, scope rules, automated validation, checklist)
|
|
||||||
|
|
||||||
## Files Affected
|
|
||||||
- `.devin/skills/dependency-graph/SKILL.md` (new)
|
|
||||||
- `.devin/skills/deployment-coordination/SKILL.md` (new)
|
|
||||||
- `.devin/skills/skill-creation/SKILL.md` (new)
|
|
||||||
- `docs/specs/GRM-173.md` (new)
|
|
||||||
|
|
||||||
## Test Plan
|
|
||||||
- Verify all three SKILL.md files follow the required structure (H1, When to Invoke, Prerequisites)
|
|
||||||
- Verify referenced make targets and file paths are accurate
|
|
||||||
- Run `make pytest-cov` to confirm no test regressions (skills are docs-only, no code changes)
|
|
||||||
- Confirm shared skills (`dependency-graph`, `skill-creation`) are ready for cross-repo sync
|
|
||||||
|
|
||||||
## Deploy Plan
|
|
||||||
- Merge to master via auto-merge workflow
|
|
||||||
- No runtime changes; documentation-only (`.devin/**` is infrastructure path, no release triggered)
|
|
||||||
|
|
||||||
## Rollback Plan
|
|
||||||
- Revert the merge commit; skill files are removed, no functional impact
|
|
||||||
|
|
||||||
## Acceptance Criteria
|
|
||||||
- [x] REQ-1: `.devin/skills/dependency-graph/SKILL.md` exists with ecosystem map, dependency chain, correct change order, and state verification sections
|
|
||||||
- [x] REQ-2: `.devin/skills/deployment-coordination/SKILL.md` exists with release flow, downstream consumer table, coordination steps, and common mistakes
|
|
||||||
- [x] REQ-3: `.devin/skills/skill-creation/SKILL.md` exists with skill structure template, quality standards, scope rules, automated validation, and creation checklist
|
|
||||||
+2
-2
@@ -36,7 +36,7 @@ ci = [
|
|||||||
"build==1.5.1",
|
"build==1.5.1",
|
||||||
"twine==6.2.0",
|
"twine==6.2.0",
|
||||||
# Reusable CI/CD and dev tools (auto-merge, pr-review, pre-push checks, etc.)
|
# Reusable CI/CD and dev tools (auto-merge, pr-review, pre-push checks, etc.)
|
||||||
"devx @ git+https://git.oblachno.oblachno.fyi/oblachno-oss/devx.git@v0.51.9",
|
"devx @ git+https://git.oblachno.oblachno.fyi/oblachno-oss/devx.git@v0.51.0",
|
||||||
]
|
]
|
||||||
# Lint and type-checking tools (validate job)
|
# Lint and type-checking tools (validate job)
|
||||||
lint = [
|
lint = [
|
||||||
@@ -56,7 +56,7 @@ molecule = [
|
|||||||
dev = [
|
dev = [
|
||||||
"grm[ci,lint,molecule]",
|
"grm[ci,lint,molecule]",
|
||||||
# Reusable CI/CD and dev tools (pre-push hooks, create-task, create-pr)
|
# Reusable CI/CD and dev tools (pre-push hooks, create-task, create-pr)
|
||||||
"devx @ git+https://git.oblachno.oblachno.fyi/oblachno-oss/devx.git@v0.51.9",
|
"devx @ git+https://git.oblachno.oblachno.fyi/oblachno-oss/devx.git@v0.51.0",
|
||||||
# Non-Python dev dependency: checkmake (Makefile linter)
|
# Non-Python dev dependency: checkmake (Makefile linter)
|
||||||
# Install via: go install github.com/checkmake/checkmake/cmd/checkmake@latest
|
# Install via: go install github.com/checkmake/checkmake/cmd/checkmake@latest
|
||||||
]
|
]
|
||||||
|
|||||||
+1
-1
@@ -1,3 +1,3 @@
|
|||||||
"""Gitea Runner Manager — lean CLI for managing Gitea Actions runners."""
|
"""Gitea Runner Manager — lean CLI for managing Gitea Actions runners."""
|
||||||
|
|
||||||
__version__ = "0.23.3"
|
__version__ = "0.22.0"
|
||||||
|
|||||||
@@ -1,120 +0,0 @@
|
|||||||
"""Pytest tests for Devin skill validation.
|
|
||||||
|
|
||||||
Validates that all skills in .devin/skills/ are well-formed: H1 title,
|
|
||||||
"when to invoke" section, prerequisites when commands are referenced,
|
|
||||||
make-target references that exist, and file references that exist.
|
|
||||||
|
|
||||||
Run with: make pytest TEST=tests/test_skills_validation.py
|
|
||||||
"""
|
|
||||||
|
|
||||||
from __future__ import annotations
|
|
||||||
|
|
||||||
import re
|
|
||||||
from pathlib import Path
|
|
||||||
|
|
||||||
import pytest
|
|
||||||
|
|
||||||
REPO_ROOT = Path(__file__).resolve().parents[2]
|
|
||||||
|
|
||||||
# Sections required for every skill
|
|
||||||
REQUIRED_SECTIONS = ["when to invoke"]
|
|
||||||
|
|
||||||
# Sections required only for skills that reference commands/tools
|
|
||||||
COMMAND_REQUIRED_SECTIONS = ["prerequisites"]
|
|
||||||
|
|
||||||
# Markers indicating a skill references commands/tools
|
|
||||||
COMMAND_MARKERS = ("`make ", "```bash", "```sh", "curl ", "python ", "python3 ", "ssh ")
|
|
||||||
|
|
||||||
EXPECTED_SKILLS = [
|
|
||||||
"dependency-graph",
|
|
||||||
"deployment-coordination",
|
|
||||||
"devx-workflow",
|
|
||||||
"molecule-testing",
|
|
||||||
"pr-review",
|
|
||||||
"runner-ops",
|
|
||||||
"skill-creation",
|
|
||||||
"spec-driven-development",
|
|
||||||
"testing-and-debugging",
|
|
||||||
"vikunja-tasks",
|
|
||||||
]
|
|
||||||
|
|
||||||
|
|
||||||
def _find_skills() -> dict[str, Path]:
|
|
||||||
skills_dir = REPO_ROOT / ".devin" / "skills"
|
|
||||||
assert skills_dir.exists(), ".devin/skills/ directory not found"
|
|
||||||
return {d.name: d / "SKILL.md" for d in skills_dir.iterdir() if d.is_dir() and (d / "SKILL.md").exists()}
|
|
||||||
|
|
||||||
|
|
||||||
# Skills shared with other repos — file-path references are only checked
|
|
||||||
# in the owning repo (infra), where the referenced files live.
|
|
||||||
SHARED_SKILLS = {"cross-repo-sync", "branch-hygiene", "dependency-graph", "skill-creation"}
|
|
||||||
|
|
||||||
|
|
||||||
def _make_targets() -> set[str]:
|
|
||||||
"""Collect make targets from Makefile plus included devx .mak files."""
|
|
||||||
targets: set[str] = set()
|
|
||||||
makefile = REPO_ROOT / "Makefile"
|
|
||||||
if makefile.exists():
|
|
||||||
targets.update(re.findall(r"^([a-zA-Z][a-zA-Z0-9_-]*):", makefile.read_text(), re.MULTILINE))
|
|
||||||
for mak in REPO_ROOT.glob(".venv/lib/python*/site-packages/devx/make/*.mak"):
|
|
||||||
targets.update(re.findall(r"^([a-zA-Z][a-zA-Z0-9_-]*):", mak.read_text(), re.MULTILINE))
|
|
||||||
return targets
|
|
||||||
|
|
||||||
|
|
||||||
def _validate_skill(skill_name: str, skill_path: Path, make_targets: set[str]) -> list[str]:
|
|
||||||
"""Validate a single skill file. Returns list of error messages."""
|
|
||||||
errors: list[str] = []
|
|
||||||
content = skill_path.read_text()
|
|
||||||
|
|
||||||
if not re.search(r"^# ", content, re.MULTILINE):
|
|
||||||
errors.append(f"{skill_name}: missing H1 title")
|
|
||||||
|
|
||||||
lower = content.lower()
|
|
||||||
for section in REQUIRED_SECTIONS:
|
|
||||||
if f"## {section}" not in lower:
|
|
||||||
errors.append(f"{skill_name}: missing '## {section.title()}' section")
|
|
||||||
|
|
||||||
references_commands = any(marker in content for marker in COMMAND_MARKERS)
|
|
||||||
if references_commands:
|
|
||||||
for section in COMMAND_REQUIRED_SECTIONS:
|
|
||||||
if f"## {section}" not in lower:
|
|
||||||
errors.append(
|
|
||||||
f"{skill_name}: missing '## {section.title()}' section "
|
|
||||||
"(required because skill references commands/tools)"
|
|
||||||
)
|
|
||||||
|
|
||||||
for target in re.findall(r"`make ([a-zA-Z][a-zA-Z0-9_-]*)`", content):
|
|
||||||
if target not in make_targets:
|
|
||||||
errors.append(f"{skill_name}: references `make {target}` but target does not exist")
|
|
||||||
|
|
||||||
# File-path checks: skip shared skills (checked in infra) and
|
|
||||||
# placeholder paths containing <...> templates.
|
|
||||||
if skill_name not in SHARED_SKILLS:
|
|
||||||
for match in re.findall(r"`((?:scripts|src|ansible|docs|tests|environments)/[^`\s]+)`", content):
|
|
||||||
if "<" in match:
|
|
||||||
continue
|
|
||||||
if not (REPO_ROOT / match).exists():
|
|
||||||
errors.append(f"{skill_name}: references `{match}` but file does not exist")
|
|
||||||
|
|
||||||
return errors
|
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize("skill_name", EXPECTED_SKILLS)
|
|
||||||
def test_skill_exists(skill_name: str) -> None:
|
|
||||||
"""Each expected skill must have a SKILL.md."""
|
|
||||||
skill = REPO_ROOT / ".devin" / "skills" / skill_name / "SKILL.md"
|
|
||||||
assert skill.exists(), f"{skill_name}/SKILL.md not found"
|
|
||||||
|
|
||||||
|
|
||||||
def test_minimum_skill_count() -> None:
|
|
||||||
"""The repo should carry a working set of skills, not a stub."""
|
|
||||||
assert len(_find_skills()) >= 8, "expected >=10 skills"
|
|
||||||
|
|
||||||
|
|
||||||
def test_all_skills_validate() -> None:
|
|
||||||
"""All skills must pass structure/reference validation."""
|
|
||||||
make_targets = _make_targets()
|
|
||||||
errors: list[str] = []
|
|
||||||
for skill_name, skill_path in _find_skills().items():
|
|
||||||
errors.extend(_validate_skill(skill_name, skill_path, make_targets))
|
|
||||||
assert not errors, "Skill validation failed:\n" + "\n".join(f" - {e}" for e in errors)
|
|
||||||
Reference in New Issue
Block a user