Digital Transformation Stories and Insights

Vandana
Vandana Unnikrishnan
9 Min Read
JIT Ordering for Produce

Explore Our Topics

Artificial Intelligence

Shadow AI and Data Leakage

Shadow AI: Building a Governance-First Security Model

GenAI is no longer an experiment in a single innovation lab. Network telemetry shows enterprise GenAI traffic surged nearly 900% in 2024. Organizations are interacting with dozens of unsanctioned GenAI apps, causing data loss incidents linked to AI to double. Today, they make up 14% of all SaaS Data Loss Prevention (DLP) events. Behind these numbers is a familiar pattern: when vetted AI tools are slowed by pay-per-token budgets or compliance controls that affect output quality, employees turn to personal subscriptions that offer fewer constraints. This is "shadow AI" in practice, and it has rapidly evolved into one of the most expensive drivers of data breaches. IBM’s 2025 Cost of a Data Breach report shows incidents involving shadow AI cost an average of $670,000 more than typical breaches, accounting for roughly 20% of all breaches today. This article outlines the active governance-first blueprint we deploy today, leveraging modern Security Service Edge (SSE) and Identity platforms, to enable rapid AI adoption while strictly containing data leakage and regulatory risk. What's in this article: How unvetted AI tools bypass standard security and inflate data breach costs Why legacy security, such as blocking domains, fails against shadow AI A four-pillar governance framework to map AI usage, enforce risk-based policies, and apply data guardrails without stifling innovation A 90-day execution blueprint to move to a secure, observable AI environment Shadow AI: How We Got Here “Shadow AI” describes AI tools and services used without IT or security teams’ approval. It can be browser plugins, SaaS copilots, GenAI APIs integrated directly by teams, and personal accounts of popular AI platforms. The intent is mostly benign: employees are trying to draft emails faster, summarize documents, debug code, or explore ideas. The problem is what goes into the prompt. Recent studies highlight the scale of this behaviour: Widespread Adoption: The average enterprise interacts with around 66 different GenAI applications, roughly 10% of which are rated high-risk. High-volume Policy violations: Organizations see an average of 223 monthly attempts to paste regulated data, source code, Intellectual Property (IP), and credentials into GenAI prompts. Inadequate Tooling: Only about half of organizations have DLP controls capable of preventing sensitive data from leaking via GenAI apps in real time. In other words, GenAI usage is exploding, sensitive data is flowing through prompts, and technical guardrails are lagging far behind. Training and awareness alone cannot close this gap. How Unvetted AI Creates Costly Data Breaches Traditional shadow IT created isolated islands of unapproved tools. Shadow AI goes further: it creates direct channels for high-value data to leave your organization entirely. Here are the key risk patterns driving this leakage: Prompt-based Exfiltration: Employees paste customer records, strategy decks, or proprietary code into prompts to "get better answers." Today, regulated data and IP make up the largest share of GenAI policy violations. For instance, an engineer pasting proprietary authentication logic into a public AI tool to debug a production outage inadvertently exposes critical IP to the vendor's training dataset. The Long Tail of GenAI Apps: Enterprises often see traffic flowing to dozens of unvetted GenAI tools. Many of these lack mature security, contractual protections, or data handling guarantees. Opaque Data Flows: Over 80% of organizations cannot see where their AI-related data ends up. They are running hundreds of unofficial apps with zero visibility. Expanded Attack Surface: OAuth integrations and AI browser extensions frequently request broad access to mailboxes, files, and code repositories. Left unchecked, they become easy pivot points for attackers. When a breach involves shadow AI, remediation is costlier and slower. Security teams are left completely blind, unsure of which tools were used, what specific data was leaked, or how long the exposure has been occurring. Why Legacy Governance Models Fail Against GenAI Most enterprises today respond to GenAI with one of three approaches: Block Everything: Outbound traffic to popular GenAI domains is blocked. This may reduce leakage in the short term, but employees often route around it via personal devices, mobile hotspots, or lesser-known tools. Awareness-first: Organizations run training, issue AI usage guidelines, and rely on user judgment. Without technical enforcement, this tends to drift back toward convenience under pressure. Tool-centric: Security teams deploy a new “AI security” platform without aligning it to existing data governance, identity, and risk processes. The underlying issue is architectural: AI is treated as a new silo instead of a channel that cuts across data, identity, cloud, and SaaS. Leading guidance from regulators and industry bodies—such as NIST’s AI Risk Management Framework (RMF) and emerging regional AI acts—stresses that AI risk must be integrated into existing governance and assurance structures, not managed in a vacuum. The Four Pillars of Governance-First GenAI Security A governance-first model for GenAI adoption can be framed around four core principles. 1. Inventory Before Innovation You cannot govern what you cannot see. Discover AI Usage: Utilize AI-powered Security Service Edge (SSE) platforms, combining Cloud Access Security Broker (CASB), endpoint DLP, and proxy capabilities, to maintain full visibility into GenAI traffic across both on-premises and remote environments. Map Data Flows: Understand which repositories (CRM, ERP, source code, file shares) are feeding prompts directly or via connectors. Catalog Tools: Establish a living register of approved AI services and explicitly track unsanctioned tools with associated risk levels. 2. Policy and Risk-Based Access Once an inventory exists, governance can move from generic warnings to concrete rules. Define AI Usage Policy: Codify what data classes (for example, payment data, health data) may never be sent to external GenAI, and under what conditions internal AI platforms may consume sensitive data. Risk-based Access Tiers: Establish rules for tools. For example, Tier 1 high-risk tools are blocked; Tier 2 moderate-risk tools allow read-only, anonymized, or synthetic data, and full usage is reserved for managed devices and within specific roles. Align with Regulations: Ensure policy explicitly references applicable frameworks and laws (NIST AI RMF, sectoral regulations, EU AI Act, etc.). 3. Guardrails in the Data Path Operationalize policy via guardrails where data flows. AI-aware DLP and CASB: Extend DLP policies to GenAI endpoints, inspecting prompts and file uploads in real time to block or redact sensitive data before it leaves the organization. Industry data shows that only around half of organizations have this in place today, despite rapidly increasing policy violations. Secure AI Gateways and Proxies: Route GenAI calls, both user and application integration traffic, through gateways that handle authentication, encryption, and logging. For example, in our architecture, integrating identity providers with SSE gateways ensures only authenticated users on compliant devices reach sanctioned GenAI tools. Pattern Libraries for Safe Prompts: Offer approved prompt templates, context windows, and integration patterns (for example, a retrieval-augmented pattern that keeps sensitive data in a private vector store). 4. Continuous Assurance and Monitoring GenAI systems evolve too quickly for a one-time risk assessment. Establish a cross-functional GenAI council to track metrics like blocked uploads and newly discovered apps. Apply risk assessment gates at each stage of the GenAI solution lifecycle, from pilot to production. Aligning AI Accountability: Three Lines of Defense Governance-first security is as much about operating model as it is about technology. A practical structure often mirrors the “three lines of defense” concept adapted for AI: Line of DefenseOwnerKey ResponsibilitiesFirst LineBusiness & Product TeamsIdentify use cases, classify data sent to AI systems, and apply approved templates for safe usage.Second LineCentral AI Governance & SecurityDefine policy/risk appetite, approve or deny GenAI tools, and operate AI gateways, DLP, and CASB controls.Third LineAudit & AssurancePeriodically review control effectiveness, policy adherence, and validate that reported usage aligns with reality. Clear Responsible, Accountable, Consulted, and Informed (RACI) Matrix definitions around GenAI use cases, from idea intake to decommissioning, are critical. Without this, shadow AI fills the vacuum as teams move faster than central functions can respond. A 90-Day Blueprint for Secure AI Adoption Enterprises often ask a simple question: Where do we start? Based on our own phased rollout and internal learnings, the following is a pragmatic 90-day plan to establish a minimum viable governance-first model. Days 1-30: Discovery Run discovery for GenAI traffic across proxies, firewalls, and CASB tools. Identify top GenAI applications in use (sanctioned and unsanctioned). Map which business units and roles are driving the highest usage. Draft an interim AI usage guideline to reduce risky behaviors while governance is being formalized. Days 31-60: Codify Policy and Implement Controls Formally sanction a dedicated GenAI governance working group. Define prohibited and restricted data categories for GenAI usage. Implement or extend DLP and CASB policies specifically for GenAI endpoints. Deploy an AI gateway or secure proxy for at least one sanctioned GenAI platform and route pilot traffic through it. Days 61-90: Operationalize and Scale Document approved GenAI usage patterns and publish them as reference architectures. Extend AI-aware DLP to additional high-risk channels (developer tools, collaboration platforms). Integrate GenAI risk checks into existing change management and solution review processes. Establish monthly metric reporting to the governance forum and quarterly summaries to executive leadership. This is not the final state, but it is sufficient to move to a controlled, observable environment where innovation and security can coexist. Characteristics of a Mature GenAI Governance Program A governance-first GenAI security model should be judged less by the number of blocked prompts and more by its ability to support safe, scalable adoption. Success indicators include: Reduced Surprise: Shadow AI tools and traffic become rare; new tools are surfaced via discovery and quickly triaged through governance. Fewer Risky Prompts: Policy violations tied to GenAI start to plateau and then decline as guardrails and training work together. Improved Breach Economics: Fewer incidents involve GenAI; when they do, logs, policies, and architectures allow faster containment. Regulatory Readiness: The organization can easily prove to auditors that its GenAI usage aligns with global AI risk management frameworks (such as the NIST AI RMF or ISO/IEC 42001) and emerging regulations (like the EU AI Act). Business Momentum: More business units confidently bring GenAI use cases to production through standardized, approved paths rather than going rogue. Ultimately, the aim is not to eliminate shadow AI by decree, but to make the sanctioned path to GenAI clearly safer, easier, and more valuable than the alternatives. Closing Thoughts Shadow AI is the predictable byproduct of powerful tools meeting unmet needs. The response cannot simply be more blocking or another point solution. Enterprises that succeed with GenAI will treat it as a first-class citizen in their governance and security architecture: visible, policy-driven, and controlled at the points where data moves. By starting with inventory, codifying risk-based policies, inserting guardrails across the data path, tracking the provenance of AI-generated assets, and anchoring everything in a clear operating model, organizations can turn GenAI from a leakage multiplier into a governed capability.
Neeraj Babu
Neeraj Babu

Quality Engineering for AI: A Framework for Validating Agent Behavior

As agentic AI takes on more decision-making and execution within enterprise applications, Quality Engineering (QE) teams face a different testing problem. Unlike traditional software, these systems are non-deterministic: the same request may follow different execution paths or generate different responses while still reaching a valid outcome. A generated response may appear correct even when the underlying agent selects the wrong API, omits critical parameters, or hallucinates information. Relying solely on the final output can create a false sense of confidence in system quality. Reliable testing of agentic systems requires QE teams to evaluate more than the final response. They need to understand how the agent reached that response, whether it made the right operational decisions, and whether those decisions were appropriate for the task. Our AI Quality Engineering Framework addresses this challenge by making intermediate agent behavior observable, measurable, and testable throughout the development lifecycle. What’s in this article: Why does agentic AI require a different approach to quality engineering? How can QE teams test the behavior of AI agents? How our AI Quality Engineering Framework evaluates execution, agent decisions, and response quality How structured evaluation helps detect regressions and support release decisions as models, prompts, and tools change Agentic AI Requires a Different QE Approach Traditional software testing typically validates whether a predefined input produces an expected outcome. With agentic AI, QE must extend that validation to the decisions and actions that lead to the outcome. This was a key consideration in our approach to testing an enterprise-grade agentic AI application for a leasing platform. The application coordinated multiple AI agents and invoked backend tools to retrieve property data dynamically. In addition to validating the application’s functional behavior, we designed our QE approach to evaluate how the agents handled each interaction. This meant testing questions that conventional functional checks do not typically address: Was the request routed to the correct specialized agent? Were the appropriate backend tools selected? Were tool parameters interpreted accurately? Was the final response grounded in retrieved data, or did the model hallucinate a plausible-sounding answer? From Observability to Behavior Evaluation Modern observability platforms such as LangSmith provide valuable visibility into execution traces, agent interactions, and tool calls. But observability alone is not enough. A trace can confirm that a tool was invoked; it cannot determine whether that tool was appropriate for the user’s intent or whether the resulting decision was correct. We, therefore, built our AI Quality Engineering Framework around a simple principle: AI behavior must be observable before it can be evaluated, and it must be evaluated before it can be trusted. AI behavior must be observable before it can be evaluated, and it must be evaluated before it can be trusted. The Three-Layer AI Evaluation Framework To move from simply observing AI behavior to actually validating it, we implemented a three-layer evaluation flow in our QE framework. Every user interaction is assessed at the execution, decision, and response-quality levels before it is considered release-ready. 1. Execution Tracing We implemented execution tracing through LangSmith instrumentation in our QE framework, so every conversation turn is captured as a structured runtime trace. For each user request, the framework records agent routing, tool calls, tool arguments, tool outputs, intermediate reasoning checkpoints, and final response metadata in sequence. Inside the framework, this works as a trace-first pipeline: test execution emits interaction events, the events are normalized into a single schema, and the run is persisted as a chronological conversation trace in LangSmith. Because traces are standardized across runs, QE can reliably replay paths, compare behaviors build-to-build, and identify exactly where the agent workflow diverged. While this layer establishes what happened during an interaction, the next layer determines whether those actions were correct. 2. Decision Validation A decision-validation layer on top of trace capture checks whether the system made the right operational choices for user intent. We configured the framework with two types of checks: Standard quality evaluators to monitor hallucination, toxicity, bias, and fairness checks Custom task evaluators tailored to business workflows For the leasing agent, we introduced a custom tool-correctness evaluator to verify that the selected backend tool, input parameters, and invocation order were appropriate for property-related questions. We added this because many failures were not hard crashes; they were decision-quality errors: the system responded fluently but used the wrong tool path or incomplete parameters, which risks inaccurate leasing information. Once we could validate the agent’s decisions, we also needed to assess the quality of the response those decisions produced. 3. LLM-as-a-Judge For response-quality assessment, we operationalized an LLM-as-a-Judge pipeline in the framework. The framework assembles inputs from three sources: Conversation context from trace history Evidence from tool outputs The agent’s final generated response The judge model scores every turn against a fixed evaluation rubric covering factual accuracy, contextual relevance, data grounding, hallucination risk, and overall usefulness, allowing quality to be tracked reliably across runs and releases. The framework aggregates these results at turn-level and conversation-level to produce comparable quality metrics over time. This makes conversational quality regression-testable: model, prompt, or toolchain changes can be measured for quality drift before release. The QE Impact: From Testing to Engineering Trust The three layers work together to give QE a more complete view of an agent interaction. For leasing conversations involving pricing, availability, and property details, failures could now be isolated to the exact layer: Incorrect agent routing Wrong tool choice Incomplete or incorrect tool parameters A mismatch between retrieved data and the generated response Low-quality final responses despite technically successful execution Essentially, our QE approach shifted from: Pass/fail conversation checks to decision-level validation Manual transcript inspection to repeatable evaluator-driven scoring “Response looks reasonable” to evidence-backed grounding and correctness checks Post-release discovery of subtle AI drift to pre-release detection in regression cycles Making AI Quality Actionable Across Teams The framework also made AI quality signals useful to different stakeholders: Quality engineers gain deterministic diagnostics and reproducible failure evidence from traces. Application engineers can identify where an agent workflow failed and which decision needs attention. Product leaders gain measurable quality signals that can inform release decisions. Clients are confident that leasing responses are correct and grounded, not just fluent. This is particularly important for enterprise AI systems where a plausible response is not sufficient evidence of correctness. Teams need a way to understand and measure the behavior behind that response. Governing AI Quality in Production Agentic AI has expanded what QE needs to validate. Functional testing remains essential, but it is only one part of assuring an AI system that can make decisions, invoke tools, and dynamically determine how to complete a task. Our AI Quality Engineering Framework addresses this by treating agent behavior as a testable engineering concern. It captures execution traces, evaluates operational decisions, and measures response quality against defined criteria. This makes AI quality more measurable across development and release cycles. Model changes, prompt changes, and toolchain changes can be evaluated against the same behavioral signals. Regressions can be identified earlier, and release decisions can be supported by evidence. For agentic systems, that is the shift QE needs to make: from asking whether the system produced an acceptable answer to evaluating whether the system behaved appropriately to produce it. Explore our Quality Engineering Services to learn how we can help you build more reliable agentic applications.
Milan
Milan Thomas
9 Min Read

GraphRAG: From Better Retrieval to Better Knowledge

Over the past two years, many enterprises have adopted Retrieval-Augmented Generation (RAG) to help AI models interact with their internal documents. For simple, single-document searches, standard RAG works very well. However, as teams try to answer more complex questions that require connecting multiple pieces of information, standard RAG often reaches its limits. GraphRAG—which combines the structured data of knowledge graphs with Large Language Models (LLMs)—is emerging as the logical next step. It allows AI to map relationships between different facts before answering a question. However, many treat GraphRAG like a simple technology update, assuming a new database will automatically yield better answers. In reality, GraphRAG is a dedicated knowledge engineering project. Organizations that see a return on their investment succeed because of clear ownership, a well-defined schema, and ongoing maintenance. What's in this article: What’s driving the shift to GraphRAG What it takes to create successful GraphRAG projects A reference architecture for production GraphRAG Our preferred frameworks and implementation approach The Drivers Behind the Shift to GraphRAG GraphRAG is moving from an experimental concept to enterprise production because of three specific shifts in the technology landscape. 1. The Limits of Similarity Search Standard vector RAG is designed to retrieve text that is semantically similar to a user's query. This is excellent for finding a specific policy, procedure, or document. However, it struggles with two types of questions common in business: Multi-hop questions: Queries that require linking facts across several different documents. For example, "Which of our vendors' suppliers are exposed to the new EU regulation?" Aggregate questions: Queries where the answer requires summarizing trends across an entire dataset, such as "What are the most common themes in this quarter's IT support tickets?" 2. Transition from Pipelines to Agents The role of enterprise AI is also changing. Today's AI agents execute workflows that involve planning, retrieving information multiple times, calling external tools, and verifying intermediate results. This requires knowledge sources that expose explicit relationships and provenance rather than isolated text chunks. GraphRAG provides that structured foundation, allowing agents to navigate connected information. 3. Automated Construction of Knowledge Graphs Until recently, building a knowledge graph required a team of human data architects to manually enter relationships, which made it too expensive for many projects. Today, LLMs can scan unstructured text and extract entities (people, companies, products) and their relationships automatically. This automation has significantly lowered the cost of building a graph. What It Takes to Build GraphRAG The success of a GraphRAG project depends largely on how the data is prepared. Knowledge Extraction at Ingestion In standard RAG, the AI has to figure out relationships at the exact moment the user asks a question. In GraphRAG, you use computing power upfront—during data ingestion—to extract relationships and store them as definitive links. When an LLM extracts "Company A supplies Company B" and saves it as a connected line in the graph, a complex question becomes a simple, fast database lookup. Ontology Design An ontology is the set of rules that defines what categories and relationships are allowed in your graph. When extracting data, it is tempting to let the AI create categories automatically. However, an open schema usually creates chaos. The AI might create WORKS_AT, EMPLOYED_BY, and WORKS_FOR as three separate relationships. Although they mean the same thing, the graph treats them as different relationships. When a user queries the graph later, the system will miss data because the relationships are fragmented. The solution is to define a strict, small schema (for example, 10 to 15 entity types) designed specifically around the questions users will ask. Entity Resolution Entity resolution is the process of ensuring that different variations of a name point to the same record. For example, "IBM", "I.B.M.", and "International Business Machines" must be merged into one single node. If they are not merged, the graph breaks into disconnected pieces, and the AI agent will hit dead ends when trying to follow a relationship. In a recent engagement with a banking fintech client, we saw this failure mode play out. Their initial attempt at open extraction yielded a noisy graph that struggled to distinguish between overlapping product features and historical software versions. We paused the build to co-design a strict, closed schema alongside their internal domain experts. That single intervention turned the system around: entity detection accuracy stabilized, and the graph could finally trace how specific product features evolved across versions—a critical business insight that was previously lost in a sea of generic, misaligned nodes. Data Provenance Agents make knowledge graphs highly effective, but they also amplify data errors. Because an agent queries the database multiple times in a loop, a single piece of bad data can compound into a major error in the final answer. Therefore, data provenance is mandatory. Every fact extracted into the graph must contain a direct link back to the source text. If an auditor or a user cannot verify where a fact came from, the system cannot be trusted. Inside the Implementation: A Reference Architecture While architectures vary by use case, most production GraphRAG systems share the same core components: a graph structure that represents enterprise knowledge, an ingestion pipeline that builds and maintains it, and retrieval patterns that combine graph traversal with LLM reasoning. The Two-Layer Graph A production GraphRAG store is not one graph but two, linked together: Lexical layer: Mirrors the source documents as Document → Section → Chunk nodes, with vector embeddings associated with each chunk. This provides the semantic retrieval capabilities of a traditional vector store within the graph. Domain layer: Represents the entities (Company, Product, Regulation) and relationships extracted from the text, with every entity connected back to the chunks it appeared in via MENTIONED_IN edges. Those cross-layer edges provide the provenance described above. Any traversal through the domain layer can end by handing the LLM the actual source text, and any suspect fact can be audited against the document it came from. The Ingestion Pipeline 1. Parse and Chunk Documents Layout-aware parsing (using tools such as Docling or Unstructured) preserves document structure-headings, tables, sections rather than reducing documents to plain text. Chunks of roughly 300–800 tokens are written into the lexical layer with their hierarchy and reading order intact, and embedded for vector search. 2. Extract Against the Schema A fast, low-cost model processes each chunk, but its output is constrained through structured output so it is impossible for the extractor to emit an entity or relationship type outside the approved ontology. The difference between putting the schema in the prompt and enforcing it at the API level is the difference between a queryable graph and fragmentation like the WORKS_AT/EMPLOYED_BY problem described earlier. A typical extraction result looks like this: { "entities": [ {"type": "Company", "name": "Acme Corp"}, {"type": "Regulation", "name": "EU Regulation 2024/17"} ], "relations": [ {"source": "Acme Corp", "type": "SUBJECT_TO", "target": "EU Regulation 2024/17", "source_chunk": "doc12#c4"} ] }3. Resolve Entities before Writing Each extracted entity is checked against the existing graph: embedding similarity generates candidate matches, and a lightweight LLM adjudicates each candidate pair ("same real-world entity, yes or no?") using both entities' descriptions and neighboring relationships as context. Confirmed matches merge into the existing node, keeping all aliases and provenance links. Because this step runs during every ingestion, new documents are incorporated into the graph incrementally without requiring a full rebuild. 4. Update Graph Incrementally Once entity resolution is complete, the graph is updated using upsert (merge) semantics, so re-processing a document never duplicates nodes or edges. Every relationship carries its source_chunk reference. 5. Build Community Summaries For corpora where users ask trend and theme questions, we run community detection (the Leiden algorithm) over the domain graph and have an LLM write hierarchical summaries of each cluster. These summaries—not raw chunks—are what answer "what are the main themes?" questions. Query Time: How the Agent Actually Uses the Graph At query time, GraphRAG works best by combining vector search with graph traversal. User queries are expressed in natural language, while graphs excel at representing explicit relationships rather than fuzzy text matches. The system therefore begins with vector search over the lexical layer to identify relevant entry points, then traverses the graph to retrieve the connected information needed to answer the question. Rather than relying on a fixed retrieval pipeline, the agent is given a small set of graph operations: search_entities(text)→ candidate entity nodes (vector + name match) get_neighbors(node, rel_type?) → adjacent nodes and relationships get_source_chunks(node) → original text behind a fact (provenance) query_graph(question) → generated Cypher, read-only, as a fallbackMost questions can be answered using the first three operations. For more complex multi-hop and aggregation queries, the agent generates a Graph query directly. The graph schema is included in the model's context, the generated query is validated before execution, and any execution errors are fed back to the model for self-correction. In practice, two or three iterations resolve most failures. The vendor-exposure example introduced earlier compiles down to a single graph traversal: MATCH (v:Company {name: "Vendor X"})(p:Product)The multi-hop question that defeats similarity search becomes one declarative query because the reasoning was already done, once, at ingestion. We expose this toolbox to AI agents over the Model Context Protocol (MCP), which has become the standard integration layer: the graph database runs behind an MCP server, and any MCP-capable agent can inspect the schema and call the tools without custom glue code. This also keeps the security controls (read-only credentials, timeouts, result limits, query logging) in one enforceable place. Frameworks We Use LayerTypical ChoicesWhen and WhyGraph databaseNeo4j (with its native vector index); or FalkorDBNeo4j for enterprise deployments: mature Cypher tooling, vector search built in, official MCP server. FalkorDB for lighter-weight or embedded deployments.Graph constructionLlamaIndex PropertyGraphIndex; the neo4j-graphrag package; LangChain LLMGraphTransformerThese handle schema-constrained extraction, embedding, and upsert plumbing. Choice depends primarily on the client's orchestration stack.Indexing pipelinesMicrosoft GraphRAG; LightRAGMicrosoft GraphRAG for aggregate and thematic questions. LightRAG when frequent updates and incremental, lower-cost indexing are more important.Agent memoryGraphiti (Zep)Best for evolving data such as user preferences or account states. Its bi-temporal model (validity intervals on edges instead of overwrites) answers both "what is true now?" and "what was true then?"Agent integrationMCP servers; native function callingMCP for portability across agent frameworks; a single place to enforce read-only access and query limits.EvaluationRAGAS plus hand-built multi-hop test setsRAGAS measures faithfulness and relevance. Multi-hop graph scenarios require domain-specific hand-crafted test cases.One selection principle worth stating explicitly: the framework choice matters far less than the decisions covered in the previous section. A team with a clean closed schema and solid entity resolution will succeed with any of these stacks; a team without them might fail with all of them. Our Implementation Approach GraphRAG technology has matured significantly faster than the enterprise's understanding of what it requires. What remains scarce is the discipline to treat enterprise knowledge as a curated, continuously maintained asset rather than a pile of documents with an index on top. At QBurst, we put those principles into practice through a phased implementation approach. We begin by identifying the business questions the system must answer before designing the ontology and knowledge model. After validating the graph within a single business domain through User Acceptance Testing (UAT), we expand incrementally to additional domains. Our Institutional Knowledge Platform and deployment accelerators reduce implementation effort while preserving this staged approach, allowing clients to move from pilot to production more quickly.
Athul Jayson
Athul Jayson
15 Min Read