Flamingo Raises $4.5M Seed Round

Back to Life at Flamingo

Building a Self-Improving AI Code Reviewer Across Repos

AIAUTOMATIONBEST PRACTICESCODE REVIEWDEVOPSWORKFLOW AUTOMATION

Aug 24, 2026

Session

Engineering

Discipline

Intermediate

Level

Michael Assraf

Michael Assraf

Founder and CEO

A walkthrough of replacing CodeRabbit with an in-house, workflow-file-driven AI code reviewer that runs weekly across multiple repos and languages. The system doesn't just review code against rules — it mines new rules from real repositories, lets tech leads approve or reject AI-suggested rules, and feeds approved rules back into its own memory so future reviews get smarter. Readers can copy the two-mode GitHub Actions setup (review mode + rule-mining mode), the rule schema (tool check / linter / AI-reasoned rule), and the cost-control tactic of running expensive whole-repo reviews on a schedule instead of on every PR.

The Workflow, Step by Step
10

  • Evaluate the off-the-shelf tool first and find its ceiling

    Tried CodeRabbit and found it worked fine for obvious, single-repo/single-class suggestions but broke down across multiple repos and multiple languages. It also required constant manual rule maintenance per repository and was more expensive than a manual alternative (pointing Claude Code directly at a PR link and asking for a review, which outperformed it).

  • Stand up two parallel generator workflows per repo

    Built a system where each repository gets two auto-generated GitHub Actions workflow files: one for "documentation" (reusing an existing, older code-documentation mechanism as a style template) and one for "code review." Each has an "Update pull request" button that opens a PR to refresh the workflow file if it's out of date.

  • Gate expensive whole-repo reviews behind PR status and a schedule

    Full-repo review ("Sweep") runs once per week across the entire repository rather than on every draft PR, because a full run can cost $20-$30 on a large repo. Per-PR review only triggers once a PR moves to "ready for review" status, not on drafts, to control cost.

  • Define what Sweep actually checks for

    The weekly full-repo run looks for cross-cutting issues a single-PR reviewer would miss: the same function duplicated in two places without being unified, violations of internal code standards, etc.

  • Build a structured rule schema before writing more rules

    Each rule can be a tool/linter check (e.g., a Pavlo-style linter rule) or an "AI rule" the model must reason about. Each rule records: source repo, source path (where it originated), an optional example of "good practice" for the AI to reference, relevant languages, exempt paths, which analyzer it applies to, and severity (blocks the PR vs. warning/info/advisory).

  • Add rules manually from trusted sources

    Took a batch of real rules sent by a colleague (Kirill) and added them into the rule store with AI assistance, verifying manually that each one actually worked before treating it as part of the "rule corpus" (the AI's persistent memory used on every future run).

  • Run rule mining as a separate mode of the same action

    The same GitHub Actions file that runs code review also runs in a "mining" mode, dispatched weekly per repository. Mining gathers reference material (e.g., CLAUDE.md-style files) plus sample files from the actual repo, then auto-suggests new candidate rules — one run produced 10 new AI-generated rule suggestions for a given repo.

  • Put a human approval gate on AI-suggested rules

    Tech leads are expected to periodically review the AI-generated rule suggestions and approve or reject each one. Approving a rule both adds it to the active rule set and further trains/improves the AI for the next mining run. A hashing check prevents the same rule from being suggested and created repeatedly.

  • Let the AI batch and open PRs from its own findings

    After the review run finishes gathering findings across a repo, the AI is prompted to group related findings into sensible clusters, then opens a separate PR per cluster (seen as 28 PRs in one run), each PR carrying a confidence level on its findings.

  • Surface the open gap around cross-repo context

    A teammate raised the case of PRs that span multiple related repositories (e.g., moving code from saslib to osslib as part of an API refactor) — the reviewer, evaluated per-repo, risks flagging deleted files as regressions when they were actually just relocated to another repo it can't see. This cross-repo/cross-tenant context gap was identified live as an open risk, not yet solved.

Tools Used and What For
5

  • CodeRabbit

    Initial off-the-shelf AI code review tool; evaluated and rejected due to weak multi-repo/multi-language performance, high maintenance burden, and cost.

  • Claude Code

    Used both informally (pointed at a raw PR link and asked to review, per a colleague's workaround) and as the underlying model powering the in-house reviewer and rule-mining runs (referred to in the transcript as "cloud/clothsono file").

  • GitHub Actions

    Two auto-generated workflow files per repo (documentation, code review); the code-review action runs in two modes — full-repo review ("Sweep," weekly) and rule mining (weekly) — dispatched per repository.

  • Custom rule-authoring UI

    Internal screen for defining tool-check, linter, or AI-reasoned rules, with fields for source repo/path, example good-practice reference, applicable languages, exempt paths, analyzer, and block-vs-warn severity.

  • Rule-mining pipeline

    Automated job that samples repo files and reference docs (CLAUDE.md-style) to auto-suggest new candidate rules for human approval, with a hash check to prevent duplicate rule suggestions.

Reusable Takeaways
7

  • Benchmark the general tool against a scoped manual workflow before building anything

    Before investing in an in-house system, they tested whether simply handing an LLM a direct link and a review prompt beat the packaged product. If a five-minute manual prompt outperforms a paid tool, that's your signal to build a thin custom layer instead of buying a black box.

  • Separate "apply known rules" from "discover new rules" as two modes of one pipeline

    Running the same automation in a review mode and a mining mode — rather than building two separate systems — keeps the rule set alive without doubling infrastructure. The mining mode's whole job is to propose, not to enforce.

  • Put a human approval gate between AI suggestions and AI memory

    Letting the AI suggest rules is cheap; letting it silently adopt them is risky. A lightweight approve/reject step by a domain expert (tech lead) before a suggestion becomes part of the model's persistent "rule corpus" prevents drift and builds trust in the system over time.

  • Throttle expensive AI runs with scope and cadence, not just prompt efficiency

    When a full-context run costs real money ($20-30 per large repo), the lever isn't a cheaper prompt — it's changing when and how often it runs (weekly sweep vs. per-PR) and gating it on state (ready-for-review vs. draft).

  • Give rules provenance, not just content

    Recording where a rule came from (source repo/path) and what "good" looks like (an example file) makes rules auditable and lets the AI ground its judgment in a concrete reference instead of an abstract instruction.

  • Have the AI cluster its own findings before creating deliverables

    Rather than opening one PR per finding, prompting the model to first group related findings into coherent PRs produces reviewable, right-sized output instead of noise.

  • Name the blast radius of your system's blind spots out loud

    The cross-repo context gap (can't tell a "deleted" file was actually relocated to a sibling repo) was surfaced as a known limitation in the room rather than papered over — worth stating explicitly so users calibrate trust correctly.

Failures and Gotchas
6

  • CodeRabbit's quality dropped off outside narrow scope

    It handled obvious, localized suggestions (single repo, single class) fine but failed to generalize across multiple repos and languages without constant manual rule maintenance per repository.

  • Off-the-shelf tool was also the more expensive option

    Manually prompting Claude Code with a PR link reportedly beat CodeRabbit's review quality for less cost — undercutting the case for the paid tool entirely.

  • Full repeated runs are genuinely expensive

    A whole-repository review run can cost $20-$30 on a large repo, which is why it's capped at once per week instead of running on every PR or every draft.

  • Duplicate rule suggestions were a real problem

    Rule mining could keep re-suggesting the same rule across runs; a hash-based check had to be added specifically to block that.

  • UI didn't always show expected data

    While demoing "files changed" confidence levels on a specific PR, the expected view didn't render as expected on that particular example ("I don't know why it doesn't show it here") — a live reminder that the tooling has rough edges.

  • Cross-repo PRs are an unsolved failure mode

    When a change spans repos (e.g., moving files from saslib to osslib), a reviewer scoped to a single repo can misread an intentional relocation as a destructive deletion, because it has no shared context linking the two PRs/repos. This was flagged live as a real worry, not yet addressed by the system.

Michael Assraf

Founder and CEO

Hey everyone, I'm Michael - founder and CEO of Flamingo. Before this, I built Vicarius, a cybersecurity company focused on vulnerability remediation, where I raised over $60M in funding. Working closely with service providers through that journey, I saw firsthand how MSPs were losing money to vendor payouts and inefficient systems - and that's when the idea for Flamingo clicked. I set out to build an open-source platform that dramatically increases MSP margins while helping them deliver better service to their clients.

Frequently Asked Questions

About OpenFrame

OpenFrame isn't built to plug into your stack. It replaces it. Instead of duct-taping a dozen tools together (RMM, MDM, SIEM, patching, remote access, each its own login and bill), we bundle it into one unified platform: RMM, MDM, monitoring, automation, remote access, patch management, security monitoring, and ticketing, plus built-in AI copilots. So "does it integrate with X?" usually means: you won't need X anymore.
Most platforms give you one piece and expect you to bolt the rest on. OpenFrame unifies the whole stack in one place, with AI copilots built in. Fewer logins, fewer bills, less duct tape.
In the cloud, on US soil. Your data stays stateside.
Both. It's built for MSPs and MSSPs alike.

MSP AI Agents

Yes. In production MSP shops today, 10% to 25% of tickets close before a human opens them. Thread alone has processed 173 million tickets across 750-plus MSP partners at 96% triage accuracy, handing back 490,000-plus technician hours. Agents own the low-risk, high-volume work (password resets, MFA enrollment, known installs, onboarding and offboarding) and flag anything that touches production data or needs judgment for a human to take.
On a five-person desk, reported deployments show $78,000 to $130,000 in annual direct labor savings, roughly 30% fewer escalations, and 15% to 20% better SLA compliance. Broader MSP adoption data adds ticket handling time cut by 45% and five to 12 points of margin, all from reclaimed capacity rather than headcount cuts.