ShopperCove
Menu
All writingBlogTopicsCategoriesAboutRSS
Blog
Categories
Observability & SRE62All categories
About

Plate 80

  1. Blog

Lily in CI: Trying Fuzz-Based Backdoor Detection on a Real Repo

Hands-on Lily (ASE 2026) on an owned C toy: rosa/lily 0.6.0 pin, directed catch of a localhost-bind trigger, clean refactor with zero flags, discovery miss, and CI cost math.

Aditya Challa·30 September 2026·8 min read

Summary
On this page
  1. Intro — what this post promises
  2. Paper claim in one paragraph
  3. What “CI-shaped” means for us
  4. Install reality (measured)
  5. Toy experiment (owned code only)
  6. Baseline (A) and refactor (A′)
  7. Toy trigger (B)
  8. Protocol
  9. What we report honestly
  10. CI cost and friction for small teams
  11. How Lily sits next to Cosign and xz checklists
  12. Verdict
  13. FAQ
  14. Further reading

​​

Intro — what this post promises

ShopperCove already summarized the ASE 2026 paper Not In My Git Yard: Catching Backdoors at Commit and Release Time in Not In My Git Yard: A Fuzzing-Based Defense…. Summaries do not fulfill “hands-on after running.”

Related links:

  • Not In My Git Yard: Catching Backdoors at Commit and Release Time
  • Not In My Git Yard: A Fuzzing-Based Defense…

This post is the practitioner attempt:

  1. Paper claim in one paragraph (cited).
  2. Install / artifact reality (GitHub binsec/rosa lily branch 0.6.0).
  3. Toy backdoor experiment on our repo only.
  4. What Lily caught, missed, or flagged noisily.
  5. CI cost / friction for a solo or small team.
  6. Verdict grounded in lab evidence.

Lab honesty (29 Sep 2026 IST): Shared Linux lab box (8 cores, ~16 GiB). Docker unavailable — built from source. Pin: rosa/lily b60403c / 0.6.0, AFL++ submodule 510ccee3 + ROSA patches, rustc 1.88.0 (system 1.85.1 failed MSRV). Owned C echo toy with a clearly labeled localhost-bind trigger. Directed campaign caught magic input (socket+bind); clean refactor 0 flags in ~3 min; discovery without magic seed missed the toy string and produced one noisy atypical on TOY_NEAR.

Ethics bar (non-negotiable):

  • Only owned or purpose-built toy repositories.
  • Defensive demonstration; no instructions for attacking third-party projects.
  • No weaponized exploit chains; a toy trigger is enough to exercise detection.
  • We did not run the repo’s sudo example submodule (out of ethics scope).

Related links:

  • Not In My Git Yard / fuzz-based defense
  • xz-utils backdoor technical deep dive
  • After xz supply-chain checklist
  • Cosign + SBOM in CI

Lily augments short CI fuzzing with system-call profile monitoring: a change is flagged only if a fuzzer input triggers behavior that is both novel versus the prior revision and atypical for the new revision’s standard profiles. We ran that loop on an owned toy.


Paper claim in one paragraph

Lily (Kokkonis, Marcozzi, Zacchiroli; ASE 2026) integrates backdoor detection into (1) CI commit vetting and (2) release vetting. While fuzzing a revised program, it monitors system-call profiles. It reports a potential backdoor only when both hold: (i) an input produces a profile absent from the revised code’s “standard” profiles (built from a regression suite evolved from prior fuzz campaigns), and (ii) the same input does not produce that behavior on the previous version. A tracer then narrows suspicious diff regions for maintainers. The paper’s evaluation (their numbers, not ours) reports high detection with low false alarms under ~10-minute AFL++ campaigns and discusses evasion + a hardened LilySelective mode against corpus poisoning.

Sources:

  • arXiv HTML
  • ASE 2026 entry
  • github.com/binsec/rosa/tree/lily

In rosa 0.6.0, Novel Lily ships as rosa-filter-diff; Atypical Lily is rosa with a phase-one corpus.


What “CI-shaped” means for us

We are not claiming to reproduce the paper’s multi-thousand CPU-hour study.

Paper settingSolo/small-team proxy (this lab)
4 cores / 16 GiB reference8-core shared lab box
~10 min fuzz windowDirected cost run ~5.5 min wall; paper budget still the planning unit
Harnessed C/C++-class PUTsTiny harnessed owned echo toy
Rosarum backdoorsToy localhost bind; no third-party implant

Install reality (measured)

StepResult
Clone binsec/rosa @lilyOK — b60403c / 0.6.0
cargo build --release on system rustc 1.85.1FAIL — deps need rustc ≥1.88
rustup 1.88.0 + rebuildOK (~25 s) — rosa, rosa-trace, rosa-filter-diff
AFL++ submodule + ROSA patches + make -jOK (~50 s); QEMU mode not built
Docker plumtrie/rosa:latestBlocked — Docker not installed
LICENSELGPL-2.1 (ROSA)

First blocking error worth keeping in the runbook: darling@0.24.1 requires rustc 1.88.0.

Config pitfall: phase-one corpus files named hello.txt made Rosa look for hello.trace. Use extensionless names (hello + hello.trace).


Toy experiment (owned code only)

Baseline (A) and refactor (A′)

Minimal stdin line protocol: ping→pong, status→ok, else echo. Refactor only renames the handler — no new syscalls.

Toy trigger (B)

Clearly labeled TOY DETECTION TARGET — NOT FOR PRODUCTION. Magic token TOY_TRIGGER_SHOPPERCOVE_ONLY opens an AF_INET socket and binds 127.0.0.1:0, then closes. Loopback only. No credentials. No remote hosts.

Protocol

  1. rosa-trace non-magic seeds on current → phase-one corpus.
  2. rosa with phase_one.corpus + AFL++ mode=standard (Atypical Lily).
  3. rosa-filter-diff against clean previous config (Novel Lily).
TrialExpected (hope)Observed (29 Sep 2026 IST)
Clean refactor A→A′ (~3.3 min)No / rare false alarm0 unique backdoors; filter-diff clean
Toy trigger A→B, magic in seeds (~42 s / ~5.5 min)Flag + retain after filter-diffCaught: is_backdoor=true; syscalls 3,41,49 (close, socket, bind); Novel Lily: “New decision: suspicious”
Toy trigger A→B, magic not in seeds (~3.3 min)May miss under short budgetMissed magic string; 1 noisy atypical on seed TOY_NEAR (empty syscall vector vs cluster [24]) — not the toy bind
Time / CPU~CI budgetLocal box $0; directed catch ≪ 10 min once seeded

Redacted decision excerpt (directed catch):

[decision]
trace_name = "main__id:000001,…,orig:magic.txt"
is_backdoor = true
reason = "syscalls"
[decision.discriminants]
trace_syscalls = [3, 41, 49]  # close, socket, bind
cluster_syscalls = []

What we report honestly

  1. Caught the toy when the magic seed was present — atypical + novel Lily agreed.
  2. Missed the toy under a short discovery budget without that seed (long secret string).
  3. One noisy atypical without a real bind — triage still required (ASAN-instrumented AFL builds may add noise).
  4. Could run after rustup + AFL++ — Docker would have been easier if available.

Paper reference rates (context only — authors’ evaluation, not ours): ~90% commit / ~83% release detection averages; ~0.2% / ~4.3% false alarms in their study.


CI cost and friction for small teams

FactorLab takeaway
Harness qualityNo harness / no __ROSA_TRACE_START ⇒ no Lily
CPU minutesDirected catch in <1 min of campaign time; plan ~10 min like the paper
Corpus hygienePhase-one must exclude the trigger; directed seed is a CI regression aid
Language / buildNative C here; Node out of scope
Maintainer UXSomeone must read decisions + filter-diff logs
ToolchainRust 1.88+, clang, patched AFL++ — or vendor Docker image

Cost sketch using GitHub Actions runner pricing (2026 Linux rates) × 10 min/commit:

Related links:

  • GitHub Actions runner pricing
Commits / weekFuzz min / commitRunner-hours / monthEst. 0.006/min)
2010~13.3 h (800 min)~$4.80
2010same~0.002/min

Lab wall minutes were free on the shared box; the table is the hosted CI projection.


How Lily sits next to Cosign and xz checklists

ControlQuestion it answers
​After xz checklist​Was the artifact swapped? Process gates?
​Cosign + SBOM CI​Who built this digest; what inventory?
Lily-style fuzz oracle (this post)Did this code change introduce triggerable abnormal behavior?

None replaces the others. xz-class social + build obfuscation still argues for layered controls — see the xz deep dive.

Related links:

  • xz deep dive

Verdict

For a harnessed C toy on our runner, Lily was runnable after rustup 1.88 + AFL++ build (Docker blocked). On toy commit B with a directed magic seed it caught the localhost bind (syscalls 3/41/49) in under a minute; Novel Lily kept the finding vs clean. Clean refactor: 0 flags in ~3 minutes. Discovery without the seed missed the toy under a short budget and produced one noisy atypical. For ShopperCove-sized teams I would pilot on critical native deps that already have fuzz harnesses, and not expect a drop-in GitHub Action tomorrow — packaging + triage are the real costs. Pair with Cosign/SBOM and the after-xz checklist.


FAQ

Q1. Will Lily have caught xz?
Do not claim certainty without a documented experiment. Layer Cosign, tarball diffs, and process controls (after-xz, xz deep dive).

Related links:

  • after-xz
  • xz deep dive

Q2. Can I drop Lily on any GitHub Action tomorrow?
Only if you already have a working fuzz harness, can build rosa/AFL++ (or use their Docker image), and can absorb CPU cost. Expect packaging work.

Q3. Is a toy backdoor unethical to publish?
A clearly labeled detection target in your own repo is standard for defensive tooling evals. Do not publish offensive how-tos against others’ software.

Q4. Lily vs Rosa?
Paper positions Rosa for clean-slate / longer binary vetting; Lily for short CI/release windows with version diffs. In 0.6.0, Novel Lily is rosa-filter-diff.

Q5. False alarms?
We saw 0 on a clean refactor in ~3 min, and 1 noisy atypical in a discovery run without the real trigger. Budget human triage.

Q6. Languages beyond C?
Follow whatever the artifact supports when you pin a commit; do not assume Node/JVM coverage.


Further reading

  • Not In My Git Yard / fuzz-based defense
  • xz-utils backdoor technical deep dive
  • After xz supply-chain checklist
  • Cosign + SBOM in CI
  • ShopperCove about
  • RSS: https://www.shoppercove.com/feed.xml
lilyrosafuzzingbackdoor detectionsupply chain securityafl++ciase 2026

Lab evidence

What I found running this

Lab 29 Sep 2026 IST. binsec/rosa lily b60403c / 0.6.0; AFL++ 510ccee3+patches; rustc 1.88.0 (1.85.1 MSRV fail); Docker blocked. Owned C toy: directed magic seed CATCH atypical+novel (syscalls 3,41,49 close/socket/bind) in <1 min campaign / ~5.5 min cost run; clean refactor 0 FA ~3.3 min; discovery without seed MISSED magic + 1 noisy atypical on TOY_NEAR. GH Actions cost sketch Linux 2-core $0.006/min → ~$4.80/mo at 20 commits×10 min. Affiliates: 0. Evidence: lab-evidence/10-lily/.

Notes when a lab post goes up

Occasional email for new hands-on reviews. No sequence and no sponsors.

Related links

  • Not In My Git Yard: A Fuzzing-Based Defense Against Commit and Release Backdoors

    A new paper proposes Lily, a tool that plugs backdoor detection into CI pipelines and release vetting workflows — aimed at the kind of supply-chain attack that's so far only been stopped by luck and manual review.

    9 Sept 2026

  • Plate 13

    Cosign + SBOM in CI: Sign and Attest a Container Image in One Workflow

    One CI shape: build an image by digest, Syft SBOM, Cosign sign + SBOM attest, verify success, and prove wrong-key/unsigned failure — the operational follow-on to after-xz.

    29 Sept 2026

  • Plate 11

    After xz: A Practical Supply-Chain Checklist for Solo and Small Teams

    Turn the xz-utils backdoor lesson into action: tarball vs git diffs, SBOM, Sigstore Cosign, SLSA provenance, and a solo/small-team release gate you can run this week.

    29 Sept 2026

On this page

  1. Intro — what this post promises
  2. Paper claim in one paragraph
  3. What “CI-shaped” means for us
  4. Install reality (measured)
  5. Toy experiment (owned code only)
  6. Baseline (A) and refactor (A′)
  7. Toy trigger (B)
  8. Protocol
  9. What we report honestly
  10. CI cost and friction for small teams
  11. How Lily sits next to Cosign and xz checklists
  12. Verdict
  13. FAQ
  14. Further reading
All writingBlogCategoriesTopicsAboutPrivacyRSS

© 2026 ShopperCove