Plate 80
Lily in CI: Trying Fuzz-Based Backdoor Detection on a Real Repo
Hands-on Lily (ASE 2026) on an owned C toy: rosa/lily 0.6.0 pin, directed catch of a localhost-bind trigger, clean refactor with zero flags, discovery miss, and CI cost math.
Aditya Challa8 min read
On this page
- Intro — what this post promises
- Paper claim in one paragraph
- What “CI-shaped” means for us
- Install reality (measured)
- Toy experiment (owned code only)
- Baseline (A) and refactor (A′)
- Toy trigger (B)
- Protocol
- What we report honestly
- CI cost and friction for small teams
- How Lily sits next to Cosign and xz checklists
- Verdict
- FAQ
- Further reading
Intro — what this post promises
ShopperCove already summarized the ASE 2026 paper Not In My Git Yard: Catching Backdoors at Commit and Release Time in Not In My Git Yard: A Fuzzing-Based Defense…. Summaries do not fulfill “hands-on after running.”
Related links:
- Not In My Git Yard: Catching Backdoors at Commit and Release Time
- Not In My Git Yard: A Fuzzing-Based Defense…
This post is the practitioner attempt:
- Paper claim in one paragraph (cited).
- Install / artifact reality (GitHub
binsec/rosalily branch 0.6.0). - Toy backdoor experiment on our repo only.
- What Lily caught, missed, or flagged noisily.
- CI cost / friction for a solo or small team.
- Verdict grounded in lab evidence.
Lab honesty (29 Sep 2026 IST): Shared Linux lab box (8 cores, ~16 GiB). Docker unavailable — built from source. Pin: rosa/lily b60403c / 0.6.0, AFL++ submodule 510ccee3 + ROSA patches, rustc 1.88.0 (system 1.85.1 failed MSRV). Owned C echo toy with a clearly labeled localhost-bind trigger. Directed campaign caught magic input (socket+bind); clean refactor 0 flags in ~3 min; discovery without magic seed missed the toy string and produced one noisy atypical on TOY_NEAR.
Ethics bar (non-negotiable):
- Only owned or purpose-built toy repositories.
- Defensive demonstration; no instructions for attacking third-party projects.
- No weaponized exploit chains; a toy trigger is enough to exercise detection.
- We did not run the repo’s sudo example submodule (out of ethics scope).
Related links:
- Not In My Git Yard / fuzz-based defense
- xz-utils backdoor technical deep dive
- After xz supply-chain checklist
- Cosign + SBOM in CI
Lily augments short CI fuzzing with system-call profile monitoring: a change is flagged only if a fuzzer input triggers behavior that is both novel versus the prior revision and atypical for the new revision’s standard profiles. We ran that loop on an owned toy.
Paper claim in one paragraph
Lily (Kokkonis, Marcozzi, Zacchiroli; ASE 2026) integrates backdoor detection into (1) CI commit vetting and (2) release vetting. While fuzzing a revised program, it monitors system-call profiles. It reports a potential backdoor only when both hold: (i) an input produces a profile absent from the revised code’s “standard” profiles (built from a regression suite evolved from prior fuzz campaigns), and (ii) the same input does not produce that behavior on the previous version. A tracer then narrows suspicious diff regions for maintainers. The paper’s evaluation (their numbers, not ours) reports high detection with low false alarms under ~10-minute AFL++ campaigns and discusses evasion + a hardened LilySelective mode against corpus poisoning.
Sources:
In rosa 0.6.0, Novel Lily ships as rosa-filter-diff; Atypical Lily is rosa with a phase-one corpus.
What “CI-shaped” means for us
We are not claiming to reproduce the paper’s multi-thousand CPU-hour study.
| Paper setting | Solo/small-team proxy (this lab) |
|---|---|
| 4 cores / 16 GiB reference | 8-core shared lab box |
| ~10 min fuzz window | Directed cost run ~5.5 min wall; paper budget still the planning unit |
| Harnessed C/C++-class PUTs | Tiny harnessed owned echo toy |
| Rosarum backdoors | Toy localhost bind; no third-party implant |
Install reality (measured)
| Step | Result |
|---|---|
Clone binsec/rosa @lily | OK — b60403c / 0.6.0 |
cargo build --release on system rustc 1.85.1 | FAIL — deps need rustc ≥1.88 |
| rustup 1.88.0 + rebuild | OK (~25 s) — rosa, rosa-trace, rosa-filter-diff |
AFL++ submodule + ROSA patches + make -j | OK (~50 s); QEMU mode not built |
Docker plumtrie/rosa:latest | Blocked — Docker not installed |
| LICENSE | LGPL-2.1 (ROSA) |
First blocking error worth keeping in the runbook: darling@0.24.1 requires rustc 1.88.0.
Config pitfall: phase-one corpus files named hello.txt made Rosa look for hello.trace. Use extensionless names (hello + hello.trace).
Toy experiment (owned code only)
Baseline (A) and refactor (A′)
Minimal stdin line protocol: ping→pong, status→ok, else echo. Refactor only renames the handler — no new syscalls.
Toy trigger (B)
Clearly labeled TOY DETECTION TARGET — NOT FOR PRODUCTION. Magic token TOY_TRIGGER_SHOPPERCOVE_ONLY opens an AF_INET socket and binds 127.0.0.1:0, then closes. Loopback only. No credentials. No remote hosts.
Protocol
rosa-tracenon-magic seeds on current → phase-one corpus.rosawithphase_one.corpus+ AFL++mode=standard(Atypical Lily).rosa-filter-diffagainst clean previous config (Novel Lily).
| Trial | Expected (hope) | Observed (29 Sep 2026 IST) |
|---|---|---|
| Clean refactor A→A′ (~3.3 min) | No / rare false alarm | 0 unique backdoors; filter-diff clean |
| Toy trigger A→B, magic in seeds (~42 s / ~5.5 min) | Flag + retain after filter-diff | Caught: is_backdoor=true; syscalls 3,41,49 (close, socket, bind); Novel Lily: “New decision: suspicious” |
| Toy trigger A→B, magic not in seeds (~3.3 min) | May miss under short budget | Missed magic string; 1 noisy atypical on seed TOY_NEAR (empty syscall vector vs cluster [24]) — not the toy bind |
| Time / CPU | ~CI budget | Local box $0; directed catch ≪ 10 min once seeded |
Redacted decision excerpt (directed catch):
What we report honestly
- Caught the toy when the magic seed was present — atypical + novel Lily agreed.
- Missed the toy under a short discovery budget without that seed (long secret string).
- One noisy atypical without a real bind — triage still required (ASAN-instrumented AFL builds may add noise).
- Could run after rustup + AFL++ — Docker would have been easier if available.
Paper reference rates (context only — authors’ evaluation, not ours): ~90% commit / ~83% release detection averages; ~0.2% / ~4.3% false alarms in their study.
CI cost and friction for small teams
| Factor | Lab takeaway |
|---|---|
| Harness quality | No harness / no __ROSA_TRACE_START ⇒ no Lily |
| CPU minutes | Directed catch in <1 min of campaign time; plan ~10 min like the paper |
| Corpus hygiene | Phase-one must exclude the trigger; directed seed is a CI regression aid |
| Language / build | Native C here; Node out of scope |
| Maintainer UX | Someone must read decisions + filter-diff logs |
| Toolchain | Rust 1.88+, clang, patched AFL++ — or vendor Docker image |
Cost sketch using GitHub Actions runner pricing (2026 Linux rates) × 10 min/commit:
Related links:
| Commits / week | Fuzz min / commit | Runner-hours / month | Est. 0.006/min) |
|---|---|---|---|
| 20 | 10 | ~13.3 h (800 min) | ~$4.80 |
| 20 | 10 | same | ~0.002/min |
Lab wall minutes were free on the shared box; the table is the hosted CI projection.
How Lily sits next to Cosign and xz checklists
| Control | Question it answers |
|---|---|
| After xz checklist | Was the artifact swapped? Process gates? |
| Cosign + SBOM CI | Who built this digest; what inventory? |
| Lily-style fuzz oracle (this post) | Did this code change introduce triggerable abnormal behavior? |
None replaces the others. xz-class social + build obfuscation still argues for layered controls — see the xz deep dive.
Related links:
Verdict
For a harnessed C toy on our runner, Lily was runnable after rustup 1.88 + AFL++ build (Docker blocked). On toy commit B with a directed magic seed it caught the localhost bind (syscalls 3/41/49) in under a minute; Novel Lily kept the finding vs clean. Clean refactor: 0 flags in ~3 minutes. Discovery without the seed missed the toy under a short budget and produced one noisy atypical. For ShopperCove-sized teams I would pilot on critical native deps that already have fuzz harnesses, and not expect a drop-in GitHub Action tomorrow — packaging + triage are the real costs. Pair with Cosign/SBOM and the after-xz checklist.
FAQ
Q1. Will Lily have caught xz?
Do not claim certainty without a documented experiment. Layer Cosign, tarball diffs, and process controls (after-xz, xz deep dive).
Related links:
Q2. Can I drop Lily on any GitHub Action tomorrow?
Only if you already have a working fuzz harness, can build rosa/AFL++ (or use their Docker image), and can absorb CPU cost. Expect packaging work.
Q3. Is a toy backdoor unethical to publish?
A clearly labeled detection target in your own repo is standard for defensive tooling evals. Do not publish offensive how-tos against others’ software.
Q4. Lily vs Rosa?
Paper positions Rosa for clean-slate / longer binary vetting; Lily for short CI/release windows with version diffs. In 0.6.0, Novel Lily is rosa-filter-diff.
Q5. False alarms?
We saw 0 on a clean refactor in ~3 min, and 1 noisy atypical in a discovery run without the real trigger. Budget human triage.
Q6. Languages beyond C?
Follow whatever the artifact supports when you pin a commit; do not assume Node/JVM coverage.
Further reading
Lab evidence
What I found running this
Lab 29 Sep 2026 IST. binsec/rosa lily b60403c / 0.6.0; AFL++ 510ccee3+patches; rustc 1.88.0 (1.85.1 MSRV fail); Docker blocked. Owned C toy: directed magic seed CATCH atypical+novel (syscalls 3,41,49 close/socket/bind) in <1 min campaign / ~5.5 min cost run; clean refactor 0 FA ~3.3 min; discovery without seed MISSED magic + 1 noisy atypical on TOY_NEAR. GH Actions cost sketch Linux 2-core $0.006/min → ~$4.80/mo at 20 commits×10 min. Affiliates: 0. Evidence: lab-evidence/10-lily/.
Related links
Not In My Git Yard: A Fuzzing-Based Defense Against Commit and Release Backdoors
A new paper proposes Lily, a tool that plugs backdoor detection into CI pipelines and release vetting workflows — aimed at the kind of supply-chain attack that's so far only been stopped by luck and manual review.
9 Sept 2026
Plate 13
Cosign + SBOM in CI: Sign and Attest a Container Image in One Workflow
One CI shape: build an image by digest, Syft SBOM, Cosign sign + SBOM attest, verify success, and prove wrong-key/unsigned failure — the operational follow-on to after-xz.
29 Sept 2026
Plate 11
After xz: A Practical Supply-Chain Checklist for Solo and Small Teams
Turn the xz-utils backdoor lesson into action: tarball vs git diffs, SBOM, Sigstore Cosign, SLSA provenance, and a solo/small-team release gate you can run this week.
29 Sept 2026