Costruire un AI Agent per la sicurezza del software: quanto costa davvero?
-
Mohamed Msaad
- 05 Oct, 2026
- 06 Mins read
L'idea, sulla carta, è estremamente interessante: prendere il codice di un'azienda, indipendentemente dal linguaggio utilizzato, e inserirlo all'interno di una pipeline capace di analizzarlo automaticamente.
L'obiettivo non sarebbe soltanto trovare vulnerabilità, ma anche individuare problemi di qualità, configurazioni rischiose, dipendenze obsolete, inefficienze, problemi architetturali e possibili miglioramenti. In prospettiva, l'AI potrebbe persino proporre una correzione e aprire automaticamente una Pull Request.
La domanda, però, è: quanto costa davvero costruire una cosa del genere?
La risposta breve è: meno di quanto si potrebbe pensare per un prototipo, molto di più per qualcosa che sia realmente affidabile in produzione.
Il problema non è creare l'AI
Il primo errore sarebbe pensare di poter costruire tutto semplicemente collegando un LLM a GitHub e chiedendogli: "analizza questo repository e trova le vulnerabilità".
Per una demo può funzionare. In un'azienda molto probabilmente no.
Un sistema serio dovrebbe combinare strumenti diversi: SAST, analisi delle dipendenze, secret scanning, container scanning, Infrastructure as Code e altri controlli già presenti nell'ecosistema DevSecOps. L'AI dovrebbe quindi stare sopra questi strumenti, non sostituirli.
Il suo vero valore sarebbe prendere centinaia di risultati differenti, comprenderne il contesto e stabilire quali siano realmente importanti.
Un semplice scanner potrebbe segnalare una vulnerabilità come "High". Un agente AI potrebbe invece analizzare il percorso del dato, le autorizzazioni, le dipendenze e l'architettura dell'applicazione e concludere che quel problema è un falso positivo — oppure che tre problemi apparentemente separati rappresentano in realtà un'unica vulnerabilità molto più grave.
Ed è qui che il progetto diventa interessante.
La complessità cresce rapidamente
Il secondo problema è la varietà.
Un'azienda reale non ha necessariamente solo applicazioni Python o Java. Può avere Java, C#, JavaScript, TypeScript, Go, C++, Rust, SQL, applicazioni mobile, Docker, Kubernetes, Terraform e configurazioni cloud.
"Supportare tutti i linguaggi" quindi non significa semplicemente aggiungere un parser per ogni linguaggio. Significa comprendere anche framework, dipendenze, build system, infrastruttura e modalità di deployment.
Per questo motivo, probabilmente non costruirei un unico agente AI.
Penserei piuttosto a un orchestratore capace di coordinare agenti specializzati: uno per la sicurezza, uno per la qualità del codice, uno per le performance, uno per l'infrastruttura e infine un agente "reviewer" incaricato di verificare i risultati prodotti dagli altri.
L'obiettivo non dovrebbe essere generare più finding possibile.
Dovrebbe essere generare finding realmente utili.
Il costo maggiore non è il modello
È facile concentrarsi sul costo delle API AI, ma probabilmente non sarà quello il problema principale. Il costo maggiore sarà lo sviluppo della piattaforma.
Un MVP credibile potrebbe richiedere 3–6 mesi con un piccolo team composto almeno da competenze backend/platform, security, AI e DevOps.
Una prima versione potrebbe limitarsi a pochi linguaggi e a una pipeline molto semplice:
Git → analisi → AI verification → risultato nella Pull Request.
Un progetto di questo tipo potrebbe richiedere indicativamente diverse migliaia di euro(si parla di cifre oltre 100k), considerando sviluppo e infrastruttura, a seconda del team e dell'ambiente aziendale.
Portarlo invece a un livello enterprise cambia completamente scala: autenticazione, RBAC, audit, gestione dei dati sensibili, multi-tenancy, observability, compliance, policy aziendali, deployment privati e supporto a decine di ecosistemi.
A quel punto si entra facilmente nell'ordine dei centinaia di migliaia di euro fino a oltre un milione, soprattutto se il prodotto deve diventare una piattaforma aziendale vera e propria.
Il punto più delicato: la fiducia
C'è poi un problema ancora più importante.
Un sistema del genere non deve soltanto trovare problemi. Deve essere credibile.
Se produce 1.000 segnalazioni e 800 sono falsi positivi, gli sviluppatori smetteranno rapidamente di utilizzarlo.
Per questo una parte fondamentale dell'architettura dovrebbe essere dedicata alla verifica dei risultati e al feedback umano.
Se uno sviluppatore indica che un finding è un falso positivo, quell'informazione può diventare parte del sistema di valutazione e contribuire a migliorare progressivamente il comportamento dell'agente.
In altre parole, il vantaggio competitivo non sarebbe necessariamente il modello AI utilizzato.
Sarebbe il patrimonio costruito nel tempo: contesto, feedback, regole, evaluation, integrazioni e conoscenza dell'ambiente aziendale.
Da dove partire?
La strada più sensata, quindi, non è provare immediatamente a supportare ogni linguaggio e ogni cloud.
Partirei da un perimetro molto più piccolo:
GitHub/GitLab → 2/3 linguaggi → SAST + dependency scanning → AI verification → commento sulla Pull Request.
Se il sistema dimostra di riuscire a trovare problemi reali, ridurre i falsi positivi e far risparmiare tempo agli sviluppatori, allora ha senso espanderlo verso container, Kubernetes, Terraform, cloud e runtime.
Il risultato finale potrebbe diventare qualcosa di molto più ambizioso di un semplice "AI code scanner": una piattaforma capace di comprendere il rapporto tra codice, dipendenze, infrastruttura e rischio, intervenendo direttamente all'interno del ciclo di sviluppo.
Ed è probabilmente questa la parte più interessante del progetto.
Non costruire un altro chatbot che legge codice, ma costruire un sistema che sappia ragionare sul software e sul rischio che quel software introduce nell'azienda.
English Version
Building an AI Agent for Software Security: What Does It Really Cost?
The idea, at least on paper, is extremely compelling: take a company’s software, regardless of the programming language used, and integrate it into a pipeline capable of automatically analyzing it.
The goal would not simply be to find vulnerabilities, but also to identify code quality issues, risky configurations, outdated dependencies, inefficiencies, architectural problems, and potential improvements. In the future, the AI could even suggest a fix and automatically open a Pull Request.
The question is: what does it actually cost to build something like this?
The short answer is: less than you might think for a prototype, but significantly more for something reliable enough to use in production.
The problem isn't creating the AI
The first mistake would be to think that you can simply connect an LLM to GitHub and ask it: "analyze this repository and find the vulnerabilities."
For a demo, it can work. In an enterprise environment, it isn't enough.
A serious system would need to combine different tools: SAST, dependency analysis, secret scanning, container scanning, Infrastructure as Code analysis, and other controls already present in the DevSecOps ecosystem.
The AI should therefore sit on top of these tools, rather than trying to replace them.
Its real value would be taking hundreds of different findings, understanding their context, and determining which ones actually matter.
A traditional scanner might report a vulnerability as "High". An AI agent could instead analyze the data flow, authorization mechanisms, dependencies, and application architecture, and determine that the issue is actually a false positive — or that three apparently unrelated findings are actually part of a single, much more serious vulnerability.
That is where the project becomes genuinely interesting.
Complexity grows very quickly
The second problem is variety.
A real company rarely has only Python or Java applications. It may have Java, C#, JavaScript, TypeScript, Go, C++, Rust, SQL, mobile applications, Docker, Kubernetes, Terraform, and multiple cloud environments.
So "supporting all languages" doesn't simply mean adding a parser for each one. It means understanding frameworks, dependencies, build systems, infrastructure, and deployment models as well.
For this reason, I probably wouldn't build a single AI agent.
I would rather build an orchestrator capable of coordinating specialized agents: one for security, one for code quality, one for performance, one for infrastructure, and finally a "reviewer" agent responsible for validating the results produced by the others.
The goal should not be to generate as many findings as possible.
It should be to generate findings that are actually useful.
The biggest cost isn't the model
It is easy to focus on the cost of AI APIs, but that probably won't be the main problem.
The biggest cost will be building the platform itself.
A credible MVP could take 3–6 months, with a small team covering at least backend/platform engineering, security, AI, and DevOps.
A first version could focus on a few languages and a very simple pipeline:
Git → analysis → AI verification → result in the Pull Request.
A project of this scope could realistically require somewhere around €100k–€300k, including development and infrastructure, depending heavily on the team and the target environment.
Taking it to enterprise-grade level changes the scale completely: authentication, RBAC, auditing, sensitive-data handling, multi-tenancy, observability, compliance, corporate policies, private deployments, and support for dozens of ecosystems.
At that point, the project can easily move from hundreds of thousands of euros to more than €1 million, especially if it is intended to become a real enterprise platform.
The most delicate part: trust
There is an even more important problem.
A system like this doesn't just need to find problems. It needs to be trusted.
If it produces 1,000 findings and 800 of them are false positives, developers will quickly stop using it.
This means that a significant part of the architecture should be dedicated to validating results and incorporating human feedback.
If a developer marks a finding as a false positive, that information can become part of the evaluation system and progressively improve the agent's behavior.
In other words, the competitive advantage wouldn't necessarily be the AI model itself.
It would be the accumulated knowledge built over time: context, feedback, rules, evaluations, integrations, and understanding of the company's environment.
Where should you start?
The most sensible approach is therefore not to immediately try to support every language and every cloud platform.
Start with a much smaller scope:
GitHub/GitLab → 2–3 languages → SAST + dependency scanning → AI verification → Pull Request comment.
If the system proves that it can identify real problems, reduce false positives, and save developers time, then it makes sense to expand into containers, Kubernetes, Terraform, cloud infrastructure, and eventually runtime analysis.
The end result could become something much more ambitious than a simple "AI code scanner": a platform capable of understanding the relationship between code, dependencies, infrastructure, and risk, directly within the software development lifecycle.
And that is probably the most interesting part of the whole project.
Don't build another chatbot that reads code.
Build a system that can reason about software and the risk that software introduces to the organization.