Flamingo Raises $4.5M Seed Round

Back to Life at Flamingo

Inside an AI-Driven Multi-Repo Dev Pipeline: Code Graph, Rules & Review

AI AGENTSAI INTEGRATIONAPIARCHITECTUREAUTOMATIONCODE REVIEWDEVOPSDOCUMENTATION

Advanced

Level

Michael Assraf

Michael Assraf

Founder and CEO

A walkthrough of "Product Hub," a set of connected tools (Code Graph, Code Rules, Code Documentation, Code Review, and Change Sets) that let a Claude-based coding skill read a ClickUp task, autonomously spin up work across dependent repos, and iterate with an AI reviewer until a PR is merge-ready. The core reusable idea: store a standing "how I work" prompt in a shared, MCP-accessible location so any engineer can pull it into a fresh AI session and get consistent, repo-aware behavior without rewriting prompts each time.

The Workflow, Step by Step
9

  • Pull a standing prompt from a shared hub instead of writing one

    Presenter keeps a saved, maintained prompt in Squawkbox (an MCP-accessible prompt hub) that tells Claude how to kick off a feature session. They just say 'please pull the latest prompt from squawkbox and execute' / 'pull feature prompt from squawkbox' instead of retyping instructions each time.

  • Skill prompts for the ClickUp task and design doc links

    After pulling the prompt, the skill asks the engineer for the feature's ClickUp task and design doc links. The engineer supplies links to the 'how I work' doc and the ClickUp task, which encode task conventions.

  • Claude's skill queries Code Graph for cross-repo dependencies

    The skill checks the live cross-repo/cross-language dependency and runtime-topology tracker (Code Graph) to detect impacted repos, e.g. osslib, Multi Platform Hub, SAS Tenant.

  • AI asks permission before creating a multi-repo Change Set

    Once dependencies are known, the skill asks permission to auto-create linked worktrees/PRs bundled into a single Change Set with a computed merge order across repos.

  • Autonomous build-and-verify loop moves the task forward

    Development proceeds without manual intervention: code gets built, verified, and the linked ClickUp task is moved to in-progress automatically.

  • Remote Code Review bot comments on the opened PR

    Code Review runs once when a PR opens, calling into Code Graph/Code Intel for context. It flags issues like missing unit tests on new functions (a static 'test rule') or duplicated functions found in other repos (a static 'duplication rule'), each with a confidence score and prompt/evidence shown.

  • Choose full-diff or incremental re-review after pushing fixes

    Reviewer can be rerun two ways: 'review the new commits' (only the diff since the PR opened) or 'review the whole diff again' (the presenter's usual choice) to re-check everything.

  • Local Claude session auto-fixes flagged review comments

    A local Claude session picks up the reviewer's PR comments via CI/CD events and iterates automatically to resolve them, looping until the PR is merge-ready.

  • Human gives thumbs up/down feedback on reviewer comments

    Engineers reply thumbs up or thumbs down to individual review comments; enough thumbs-down lowers that rule's confidence score over time, tuning the reviewer's reliability.

Tools Used and What For
8

  • Product Hub (MCP server)

    Central hub exposing Code Graph, Code Rules, Code Documentation, Code Review, and Change Sets to any engineer's AI session via MCP.

  • Code Graph

    Live, cross-repo/cross-language dependency and runtime-topology tracker (scrapes Helm charts, Terraform, tracks files/symbols) used by the reviewer and doc generator for context and to catch configuration drift.

  • Code Rules

    Weekly job that mines candidate lint/AI rules from the codebase; tech leads approve or decline each rule (set as error/warning/advisory, blocking or advisory) before it feeds the reviewer's 'memory'.

  • Code Review (replacing CodeRabbit)

    Change-set-aware AI PR reviewer; checks for duplication against a teammate's branch/overlay changes rather than just main, runs on PR open, supports blocking vs advisory mode and configurable models/custom instructions.

  • Code Documentation

    Runs across a repository to build the docs folder, main README, copy licenses, and trigger full graph builds; deployed via a GitHub Actions workflow, runs weekly or on demand (expensive to run).

  • Claude (coding skill)

    Reads the ClickUp task and standing prompt, queries Code Graph, proposes and builds the multi-repo Change Set, and later auto-fixes reviewer-flagged issues in a local session.

  • Squawkbox

    Prompt-hub MCP server where the presenter stores their standing 'how I work' feature-kickoff prompt for reuse across sessions.

  • ClickUp

    Task source; the AI pulls the feature task (and linked design docs) from here to start a session and moves it to in-progress automatically as work proceeds.

Reusable Takeaways
7

  • Store a standing prompt in a shared, tool-accessible hub

    Instead of rewriting kickoff instructions every session, save a maintained prompt somewhere any AI session can fetch it (MCP server, internal wiki with an API, etc.) so behavior stays consistent across engineers and over time.

  • Give the AI a live dependency graph, not just local file context

    A cross-repo/cross-language dependency and runtime map lets an AI reviewer or agent catch issues (duplication, breaking changes) that are invisible from a single repo's diff.

  • Mine rules from real code, but gate them with human approval

    Auto-generating lint/AI rules from the codebase surfaces patterns humans wouldn't think to write by hand, but each rule needs a human (e.g. tech lead) to approve, decline, and set severity before it's trusted.

  • Attach confidence scores and feedback loops to AI-generated rules

    Let humans thumbs up/down individual AI comments or rules so unreliable ones automatically lose confidence over time instead of silently annoying everyone forever.

  • Let the agent ask permission before taking multi-system actions

    When an AI agent is about to fan out changes across multiple repos/PRs, have it propose the plan (e.g. a bundled Change Set with computed merge order) and get a go-ahead rather than acting unilaterally.

  • Support both incremental and full re-checks after fixes

    Offer a 'review just what changed' mode and a 'review everything again' mode so reviewers can pick the right tradeoff between speed and thoroughness after pushing fixes.

  • Close the loop: reviewer comments feed an auto-fix agent

    Wiring CI/CD events so a coding agent automatically picks up reviewer feedback and iterates removes the manual copy-paste-fix cycle between review tools and coding sessions.

Failures and Gotchas
5

  • AI-mined rules need human vetting before they're trusted

    Rules mined weekly from the codebase are suggestions only; a tech lead must approve or decline each one, since AI-based rule mining includes false or low-value candidates.

  • Rule confidence can degrade from negative feedback

    If a rule/comment gets a lot of thumbs-down from engineers, its confidence score drops over time — meaning some AI review comments are known to be unreliable until tuned.

  • Full documentation runs are expensive

    Code Documentation is described as 'very expensive' to run, which is why it's scheduled weekly or run on demand rather than on every change.

  • CodeRabbit was replaced due to weak multi-repo support

    The team's prior AI review tool struggled with multi-repo and graph-aware scanning, which motivated building an in-house change-set-aware Code Review tool instead.

  • Early-adoption friction expected

    The team's forward goal (merging to main quickly behind feature flags with full regression/QA on PR environments) is still new, and the presenter explicitly expects friction as adoption ramps up.

Take This With You

Michael Assraf

Founder and CEO

Hey everyone, I'm Michael - founder and CEO of Flamingo. Before this, I built Vicarius, a cybersecurity company focused on vulnerability remediation, where I raised over $60M in funding. Working closely with service providers through that journey, I saw firsthand how MSPs were losing money to vendor payouts and inefficient systems - and that's when the idea for Flamingo clicked. I set out to build an open-source platform that dramatically increases MSP margins while helping them deliver better service to their clients.

Frequently Asked Questions

About OpenFrame

OpenFrame isn't built to plug into your stack. It replaces it. Instead of duct-taping a dozen tools together (RMM, MDM, SIEM, patching, remote access, each its own login and bill), we bundle it into one unified platform: RMM, MDM, monitoring, automation, remote access, patch management, security monitoring, and ticketing, plus built-in AI copilots. So "does it integrate with X?" usually means: you won't need X anymore.
Most platforms give you one piece and expect you to bolt the rest on. OpenFrame unifies the whole stack in one place, with AI copilots built in. Fewer logins, fewer bills, less duct tape.
In the cloud, on US soil. Your data stays stateside.
Both. It's built for MSPs and MSSPs alike.

MSP AI Agents

Yes. In production MSP shops today, 10% to 25% of tickets close before a human opens them. Thread alone has processed 173 million tickets across 750-plus MSP partners at 96% triage accuracy, handing back 490,000-plus technician hours. Agents own the low-risk, high-volume work (password resets, MFA enrollment, known installs, onboarding and offboarding) and flag anything that touches production data or needs judgment for a human to take.
On a five-person desk, reported deployments show $78,000 to $130,000 in annual direct labor savings, roughly 30% fewer escalations, and 15% to 20% better SLA compliance. Broader MSP adoption data adds ticket handling time cut by 45% and five to 12 points of margin, all from reclaimed capacity rather than headcount cuts.