A walkthrough of replacing CodeRabbit with an in-house, workflow-file-driven AI code reviewer that runs weekly across multiple repos and languages. The system doesn't just review code against rules — it mines new rules from real repositories, lets tech leads approve or reject AI-suggested rules, and feeds approved rules back into its own memory so future reviews get smarter. Readers can copy the two-mode GitHub Actions setup (review mode + rule-mining mode), the rule schema (tool check / linter / AI-reasoned rule), and the cost-control tactic of running expensive whole-repo reviews on a schedule instead of on every PR.
Building a Self-Improving AI Code Reviewer Across Repos
The Workflow, Step by Step10
Evaluate the off-the-shelf tool first and find its ceiling
Tried CodeRabbit and found it worked fine for obvious, single-repo/single-class suggestions but broke down across multiple repos and multiple languages. It also required constant manual rule maintenance per repository and was more expensive than a manual alternative (pointing Claude Code directly at a PR link and asking for a review, which outperformed it).
Stand up two parallel generator workflows per repo
Built a system where each repository gets two auto-generated GitHub Actions workflow files: one for "documentation" (reusing an existing, older code-documentation mechanism as a style template) and one for "code review." Each has an "Update pull request" button that opens a PR to refresh the workflow file if it's out of date.
Gate expensive whole-repo reviews behind PR status and a schedule
Full-repo review ("Sweep") runs once per week across the entire repository rather than on every draft PR, because a full run can cost $20-$30 on a large repo. Per-PR review only triggers once a PR moves to "ready for review" status, not on drafts, to control cost.
Define what Sweep actually checks for
The weekly full-repo run looks for cross-cutting issues a single-PR reviewer would miss: the same function duplicated in two places without being unified, violations of internal code standards, etc.
Build a structured rule schema before writing more rules
Each rule can be a tool/linter check (e.g., a Pavlo-style linter rule) or an "AI rule" the model must reason about. Each rule records: source repo, source path (where it originated), an optional example of "good practice" for the AI to reference, relevant languages, exempt paths, which analyzer it applies to, and severity (blocks the PR vs. warning/info/advisory).
Add rules manually from trusted sources
Took a batch of real rules sent by a colleague (Kirill) and added them into the rule store with AI assistance, verifying manually that each one actually worked before treating it as part of the "rule corpus" (the AI's persistent memory used on every future run).
Run rule mining as a separate mode of the same action
The same GitHub Actions file that runs code review also runs in a "mining" mode, dispatched weekly per repository. Mining gathers reference material (e.g., CLAUDE.md-style files) plus sample files from the actual repo, then auto-suggests new candidate rules — one run produced 10 new AI-generated rule suggestions for a given repo.
Put a human approval gate on AI-suggested rules
Tech leads are expected to periodically review the AI-generated rule suggestions and approve or reject each one. Approving a rule both adds it to the active rule set and further trains/improves the AI for the next mining run. A hashing check prevents the same rule from being suggested and created repeatedly.
Let the AI batch and open PRs from its own findings
After the review run finishes gathering findings across a repo, the AI is prompted to group related findings into sensible clusters, then opens a separate PR per cluster (seen as 28 PRs in one run), each PR carrying a confidence level on its findings.
Surface the open gap around cross-repo context
A teammate raised the case of PRs that span multiple related repositories (e.g., moving code from saslib to osslib as part of an API refactor) — the reviewer, evaluated per-repo, risks flagging deleted files as regressions when they were actually just relocated to another repo it can't see. This cross-repo/cross-tenant context gap was identified live as an open risk, not yet solved.
Tools Used and What For5
CodeRabbit
Initial off-the-shelf AI code review tool; evaluated and rejected due to weak multi-repo/multi-language performance, high maintenance burden, and cost.
Claude Code
Used both informally (pointed at a raw PR link and asked to review, per a colleague's workaround) and as the underlying model powering the in-house reviewer and rule-mining runs (referred to in the transcript as "cloud/clothsono file").
GitHub Actions
Two auto-generated workflow files per repo (documentation, code review); the code-review action runs in two modes — full-repo review ("Sweep," weekly) and rule mining (weekly) — dispatched per repository.
Custom rule-authoring UI
Internal screen for defining tool-check, linter, or AI-reasoned rules, with fields for source repo/path, example good-practice reference, applicable languages, exempt paths, analyzer, and block-vs-warn severity.
Rule-mining pipeline
Automated job that samples repo files and reference docs (CLAUDE.md-style) to auto-suggest new candidate rules for human approval, with a hash check to prevent duplicate rule suggestions.
Reusable Takeaways7
Benchmark the general tool against a scoped manual workflow before building anything
Before investing in an in-house system, they tested whether simply handing an LLM a direct link and a review prompt beat the packaged product. If a five-minute manual prompt outperforms a paid tool, that's your signal to build a thin custom layer instead of buying a black box.
Separate "apply known rules" from "discover new rules" as two modes of one pipeline
Running the same automation in a review mode and a mining mode — rather than building two separate systems — keeps the rule set alive without doubling infrastructure. The mining mode's whole job is to propose, not to enforce.
Put a human approval gate between AI suggestions and AI memory
Letting the AI suggest rules is cheap; letting it silently adopt them is risky. A lightweight approve/reject step by a domain expert (tech lead) before a suggestion becomes part of the model's persistent "rule corpus" prevents drift and builds trust in the system over time.
Throttle expensive AI runs with scope and cadence, not just prompt efficiency
When a full-context run costs real money ($20-30 per large repo), the lever isn't a cheaper prompt — it's changing when and how often it runs (weekly sweep vs. per-PR) and gating it on state (ready-for-review vs. draft).
Give rules provenance, not just content
Recording where a rule came from (source repo/path) and what "good" looks like (an example file) makes rules auditable and lets the AI ground its judgment in a concrete reference instead of an abstract instruction.
Have the AI cluster its own findings before creating deliverables
Rather than opening one PR per finding, prompting the model to first group related findings into coherent PRs produces reviewable, right-sized output instead of noise.
Name the blast radius of your system's blind spots out loud
The cross-repo context gap (can't tell a "deleted" file was actually relocated to a sibling repo) was surfaced as a known limitation in the room rather than papered over — worth stating explicitly so users calibrate trust correctly.
Failures and Gotchas6
CodeRabbit's quality dropped off outside narrow scope
It handled obvious, localized suggestions (single repo, single class) fine but failed to generalize across multiple repos and languages without constant manual rule maintenance per repository.
Off-the-shelf tool was also the more expensive option
Manually prompting Claude Code with a PR link reportedly beat CodeRabbit's review quality for less cost — undercutting the case for the paid tool entirely.
Full repeated runs are genuinely expensive
A whole-repository review run can cost $20-$30 on a large repo, which is why it's capped at once per week instead of running on every PR or every draft.
Duplicate rule suggestions were a real problem
Rule mining could keep re-suggesting the same rule across runs; a hash-based check had to be added specifically to block that.
UI didn't always show expected data
While demoing "files changed" confidence levels on a specific PR, the expected view didn't render as expected on that particular example ("I don't know why it doesn't show it here") — a live reminder that the tooling has rough edges.
Cross-repo PRs are an unsolved failure mode
When a change spans repos (e.g., moving files from saslib to osslib), a reviewer scoped to a single repo can misread an intentional relocation as a destructive deletion, because it has no shared context linking the two PRs/repos. This was flagged live as a real worry, not yet addressed by the system.

Founder and CEO
Hey everyone, I'm Michael - founder and CEO of Flamingo. Before this, I built Vicarius, a cybersecurity company focused on vulnerability remediation, where I raised over $60M in funding. Working closely with service providers through that journey, I saw firsthand how MSPs were losing money to vendor payouts and inefficient systems - and that's when the idea for Flamingo clicked. I set out to build an open-source platform that dramatically increases MSP margins while helping them deliver better service to their clients.