Live webinar The Wrong Security Hire Burns Your B2B GTM Pipeline. A fireside chat for founders · Oct 6, 9:30am PTThe wrong security hire · Oct 6 Save your seat

Blog

13 min read Updated September 6, 2026

Neuro-Symbolic AI: Why Enterprises Need More Than Large Language Models

LLMs handle the fuzzy parts: language, perception, pattern. Symbolic systems handle the parts that must be correct. The enterprises getting AI right are composing both. Here is what that stack looks like: where each half is strong, how the routing pattern works, and which tasks belong to which.

Neuro-symbolic AI combines two families of technique. The neural side uses deep learning and large language models for perception, language, and pattern recognition. The symbolic side uses knowledge graphs, logic rules, and classical algorithms for structured reasoning, rigorous calculation, and explainable decisions. The short version: LLMs handle the fuzzy parts, symbolic systems handle the parts that must be correct.

The approach isn’t new. Symbolic AI dominated from the 1960s through the 1980s before the neural revolution displaced it. What changed is that the neural side matured enough to carry its weight, and researchers proved the neural side alone can’t solve every enterprise problem. On episode 93, Jacob Andra, CEO of Talbot West, framed it: “LLMs are not the whole of AI. They are one tool. You build real enterprise AI by composing multiple types of AI, each doing what it is good at.”

This post covers the technology half of that conversation: what the two families are, where each is strong, how the enterprise routing pattern fits them together, and which tasks belong to which. The other half, the scoping discipline that decides what to build at all, is in Why 95% of Enterprise AI Projects Fail.

What is neuro-symbolic AI? The two-sentence version

Neuro-symbolic AI is a system design in which neural networks (today, mostly LLMs) handle perception and language while symbolic components (knowledge graphs, rules engines, solvers, classical algorithms) handle the reasoning, the arithmetic and the decisions. The neural side is probabilistic and learns from data; the symbolic side is deterministic and works from explicit representations a person can inspect.

IBM Research, which has run a neuro-symbolic program for years, describes the goal as “augmenting and combining the strengths of statistical AI, like machine learning, with the capabilities of human-like symbolic knowledge and reasoning” (IBM Research, Neuro-symbolic AI). DARPA tells the same story as history. Its AI Next campaign describes a first wave of AI built on “handcrafted knowledge,” the expert systems that “captured the specialized knowledge of experts in rules”; a second wave of machine learning that “applies statistical and probabilistic methods to large data sets”; and a third wave it is funding now, “machines that understand and reason in context.” Each wave has a weakness the other covers. DARPA’s summary of the first: “the need to handcraft rules is costly and time-consuming.” The second wave’s weakness is the subject of the next two sections.

The analogy researchers keep reaching for is Daniel Kahneman’s two systems of thinking. System 1 is fast, intuitive, pattern-matching, and occasionally confidently wrong; System 2 is slow, deliberate, and shows its work. An LLM is a very good System 1. Your pricing engine is System 2. Nobody would let System 1 file their taxes, and nobody would ask System 2 to read a customer’s angry email and work out what they actually want.

Two YSecurity team members talking over coffee away from the Black Hat USA 2026 conference floor
Two of the team over coffee at Black Hat USA 2026. Episode 93 with Jacob Andra and Stephen Karafiath of Talbot West is linked at the end of this post.

Where are large language models strong, and where do they hit a wall?

Give an LLM credit where it is due. It reads messy human input, ambiguity included, better than any software before it. It drafts, summarizes, translates and searches across enormous corpora. It generalizes to inputs nobody anticipated, which is precisely what rules cannot do. In retrieval, support, and developer productivity, that has unlocked real value, and the people who dismiss it are as wrong as the people who think it is the whole of AI.

The failure modes are just as consistent, and they show up the moment the model is the only AI layer in a workflow. It hallucinates on factual questions. It reasons inconsistently: ask the same question twice with a small rewording and get different answers. It cannot explain a decision in terms a regulator will accept, because a plausible paragraph about why is not the same thing as the actual reason. And it approximates arithmetic and constraints instead of satisfying them.

“When the result has to be right, and the result has to be explainable, an LLM by itself will disappoint you.”

Stephen Karafiath, technical co-founder of Talbot West and a 30-year veteran of GE and Oracle, on episode 93

Comparison matrix of neural AI (large language models) versus symbolic AI (knowledge graphs, rules engines, solvers) across seven rows: LLMs are strong on language and ambiguity, recall and summarization, novel unstructured inputs and learning from data; symbolic systems are strong on determinism, auditability and hard constraints and arithmetic, but their rules and graphs must be handcrafted
Where each half is strong. The filled squares never line up in one column, which is the whole argument for using both.

Security teams recognize the pattern from agentic deployments, where LLM-driven agents take inconsistent actions and leave no clean trail of why. Same model, same failure mode, higher stakes: an agent that decides its own permissions on the fly is a System 1 doing System 2’s job.

Why won’t a bigger model fix hallucination?

A common response is “we just need a bigger model.” The math says otherwise, and it says so in two papers worth knowing by name.

Ziwei Xu, Sanjay Jain and Mohan Kankanhalli’s paper Hallucination is Inevitable: An Innate Limitation of Large Language Models (first posted January 2024, revised February 2025) uses results from learning theory to show that “LLMs cannot learn all the computable functions and will therefore inevitably hallucinate if used as general problem solvers.” The formal world they prove it in is simpler than any real deployment, and they argue the real world, being more complicated, inherits the result. A broader survey, On the Fundamental Limits of LLMs at Scale (Mohsin and colleagues, November 2025), works through five limits of scaling, hallucination first, and concludes that hallucination “reflects intrinsic computational and statistical limits that no (present) architecture, scale, or optimization can overcome.” Scaling improves average quality. It cannot drive the error to zero.

“The belief that a large enough LLM will solve reasoning is a comforting story. The math says otherwise.”

Jacob Andra, CEO of Talbot West, on episode 93

Read the right way, this is good news. It tells you exactly where to stop spending: not on the next model tier, but on the structure around the model. For a general problem solver, some residual error is a law of nature. For a system that looks a fact up in a graph and applies a rule, the error on that step is zero, and the LLM’s error is confined to the wording.

Prefer a whiteboard? IBM Technology walks through the neural half, the symbolic half and why the combination matters, in six minutes, with roughly 30,000 views. Watch on YouTube.

How do knowledge graphs and rules engines fill in what LLMs can’t?

Three symbolic components do most of the work in enterprise systems, and none of them is exotic.

A knowledge graph stores facts as entities and typed relationships: this customer holds this contract, this contract is governed by this policy version, this policy supersedes that one as of this date. Every fact carries provenance, where it came from and when. The point is that the system looks facts up instead of asking a model to remember them, and a surprising share of “hallucination” problems turn out to be “no grounding data” problems in disguise. Building the graph is also where your data discipline shows: you cannot encode a fact you cannot find, which is the argument for an AI data bill of materials before the modelling starts.

A rules engine holds the business logic as explicit, version-controlled, testable if-then statements: who is eligible, what the price is, which approvals a change needs, which access a role gets. The same rule fires the same way every time, and you can print the rule that fired. Solvers and optimizers extend this to problems with hard constraints (scheduling, allocation, routing), and probabilistic models handle risk scoring where the assumptions should be written down rather than learned in the dark. Rules engines and statistical scoring models have made credit and fraud decisions for decades; what is new is a front end that can read the investigator’s question in plain English.

The verifier is the piece most teams forget. Put a deterministic check after the model and the LLM becomes a proposer while the symbolic layer becomes the judge. The clearest example is code: LLMs write it, and tests, static analysis and policy checks decide whether it merges, which is the pattern behind vibe coding security and the review gates in a secure AI SDLC.

There is a whole taxonomy of ways to wire the two halves together (Henry Kautz’s AAAI lecture on the third AI summer counts six), but two shapes cover most enterprise work: a symbolic system that calls the model for perception and language, the way AlphaGo’s tree search called a neural network to evaluate positions, and a model that calls symbolic tools for facts and decisions. The routing pattern below is the second, with the first tucked inside it.

The enterprise pattern: the LLM parses, the symbolic layer decides, the LLM explains

A neuro-symbolic system uses the LLM where it is strong and hands off where it is weak. A typical enterprise pattern has four steps:

  1. An LLM parses the question and routes it. It reads the plain-language request, extracts the intent, the entities and the constraints, and decides which system should handle it.
  2. A knowledge graph or database returns the facts. Each fact arrives with its source and its date.
  3. A rules engine, probabilistic model or solver applies the business logic. Deterministically: same input, same output.
  4. The LLM explains the result in plain language, citing the facts it used and the rule that fired.
Diagram of the enterprise neuro-symbolic AI pattern in four stages: an LLM parses and routes the plain-language question, a knowledge graph returns the facts with sources and dates, a rules engine or solver applies the business logic deterministically, and an LLM writes the answer citing the source and the rule; the audit trail records the source named, the rule named and a reproducible result
The enterprise pattern. The filled boxes are the parts that must be correct; the outlined boxes are the parts that must be understood.

Every step leaves an audit trail. The facts came from a specific source; the rule applied is named and versioned; the numeric result is reproducible. The LLM is responsible only for the language, which it is genuinely good at, and a wrong word in an explanation is cheap in a way a wrong number in a decision never is.

Amazon’s own account of Rufus, its shopping assistant, shows the same instinct at consumer scale, even though Amazon never uses the phrase neuro-symbolic. Rufus is a custom LLM, but before it answers, it “pulls information from sources it knows to be reliable, such as customer reviews, the product catalogue, and community questions and answers, along with calling relevant Stores APIs” (Amazon Science, October 2024). The model is not trusted to remember the catalogue. It is trusted to understand the shopper and to phrase the answer, and the structured systems supply the facts. The same shape appears in identity: Deepfake Attacks Keep Working Because We Keep Detecting Instead of Proving makes the case that a cryptographic identity check is a deterministic decision, a deepfake detector is a classifier, and the deterministic check should be in charge.

Why the audit trail is the feature enterprises are actually buying

Enterprise AI committees, auditors and regulators ask three questions about any automated decision: how was it made, what data did it use, and can you reproduce it. A free-text explanation from a model answers none of them reliably. A neuro-symbolic decision answers all three by construction, which is why the architecture conversation and the governance conversation are the same conversation. It is the evidence an AI management system under ISO 42001 exists to produce, and it is the difference between an AI access control audit that reads a policy file and one that has to guess what a model would have allowed.

The research community is only now catching up with that priority. A PRISMA systematic review of 167 neuro-symbolic papers published from 2020 to 2024 found 63% concentrated on learning and inference and 28% on explainability and trustworthiness, with meta-cognition, a system’s ability to monitor its own reasoning, at 5%. The practitioners are ahead of the papers here, and the enterprises paying for auditable decisions are the reason.

Which tasks belong to an LLM, and which need symbolic AI?

Two questions sort most workloads.

Does a wrong answer cost real money, safety, or a regulator’s trust? If not, an LLM on its own is fine, with a person reviewing before anything irreversible happens: drafting, summarizing, search, support triage, brainstorming, first-pass classification. Language is the whole job, and the model is the best tool for that job anyone has ever had.

If it does, can the rule, the policy or the math be written down and tested? If yes, the symbolic layer decides and the LLM is the interface: pricing, eligibility, approvals, access decisions, compliance checks, calculations, safety interlocks. If the logic genuinely cannot be written down (a judgement call about a novel document, say), ground the model in facts you control, put a checker after it, and keep a person signing off on the high-stakes outputs.

Decision guide for choosing between an LLM and symbolic AI: if a wrong answer is cheap, an LLM on its own is fine with a person reviewing; if a wrong answer is expensive and the rule can be written down and tested, the symbolic layer decides and the LLM is the interface; if the rule cannot be written down, ground the LLM in your own facts, add a checker and keep a person signing off
Two questions, three answers. Most real workflows contain all three, one step at a time.
TaskApproachWhy
Summarize a support ticket and draft a replyLLM alone; a person sends itLanguage is the job, and a wrong word is cheap
Decide whether a customer qualifies for a refundRules engine decides; LLM explainsThe policy exists and the decision must be reproducible
Grant an agent access to a production systemPolicy engine decides; never the modelAuthorization has to be deterministic and logged
Answer “which contracts renew next quarter?”LLM parses; knowledge graph answers; LLM citesThe facts live in a system, not in the model’s memory
Generate a database migrationLLM writes; tests and static analysis judgeProposer and judge
Score a transaction for fraudScoring model and rules decide; LLM writes the case noteDecades of deterministic practice, with a new interface

Most real workflows contain all three categories, and the work of deciding which step is which, before anyone picks a model, is the discipline Jacob and Stephen describe on the episode, mapping the whole organization to find the right entry point. It is also the subject of the sibling post on scoping. The AI readiness pre-flight check is the same idea from the buyer’s side: before you sign for a system that needs a knowledge graph, confirm the data to fill it exists.

For the long version, MIT's Introduction to Deep Learning course (6.S191) has a 41-minute lecture on neurosymbolic AI by Alexander Amini, with roughly 95,000 views. Worth an evening if you are going to build one of these. Watch on YouTube.

How do you add symbolic capability without starting over?

You don’t have to rip out existing LLM investments. Four moves add symbolic capability where it matters:

  1. Identify the failure surface. List where the current stack produces wrong or unexplainable answers, and what each one costs. That list is your priority order.
  2. Add a knowledge graph or structured data layer under the highest-cost failures. Many “hallucination” problems are “no grounding data” problems, and they disappear the day the model can look the fact up.
  3. Introduce a symbolic verifier after the LLM for high-stakes outputs, turning the LLM into a proposer and the symbolic layer into a judge.
  4. Keep the LLM in the interface layer, where natural language is its real advantage, and move decisions out of it one at a time.

Two notes from the security side of our practice. First, the highest-value place for a symbolic layer is anywhere the AI touches permissions. An agent that asks a model whether it is allowed to do something has a non-deterministic authorization check, which is a contradiction in terms; the answer should come from a policy engine, and the agent should have declared what it needs before it ever ran. Second, be honest about which half you are building. Wrapping a model API in a workflow is a software project with software timelines. Building the symbolic core, a graph that is actually correct and rules that actually match the business, is closer to deep tech: it has to be proven before it can be sold, and the proof is the moat.

The decision: where does the answer have to be right?

Every enterprise AI design reduces to one question per step of the workflow: does this answer have to be right, or does it only have to be useful? Useful belongs to the LLM. Right belongs to a system that can show its work, with the LLM translating at both ends. The companies getting this right did not buy a bigger model. They drew that line.

“If you cannot explain why the answer is right, and you cannot reproduce it, you are not ready to ship it. The tool for that is not a bigger LLM. The tool is a stack that knows what each part is doing.”

Jacob Andra, CEO of Talbot West, on episode 93

If your enterprise buyers are already asking how your AI makes its decisions, an ISO 42001 readiness conversation will show you how much of that evidence your architecture already produces and where a symbolic layer would produce the rest. And if you want the whole argument from the people who make it for a living, Jacob and Stephen spend 45 minutes on it in episode 93, including which AI technologies they think are overrated and which are quietly underrated.

Neuro-symbolic AI frequently asked questions

What is neuro-symbolic AI?
Neuro-symbolic AI combines neural networks, including large language models, with symbolic systems such as knowledge graphs, logic rules and solvers. The neural side handles perception, language and pattern recognition; the symbolic side handles structured reasoning, exact calculation and explainable decisions. IBM Research describes it as combining the strengths of statistical AI with human-like symbolic knowledge and reasoning. In an enterprise system that usually means an LLM at the interface and a deterministic layer making the decisions that must be correct.
What is the difference between neural and symbolic AI?
Neural AI learns patterns from data and produces probabilistic outputs, which makes it strong on language, ambiguity and inputs it has never seen, and weak on guarantees. Symbolic AI works from explicit representations such as rules, ontologies and knowledge graphs, so the same input always produces the same output and the reasoning can be printed. Symbolic knowledge has to be written down by people, which is its cost. The two are complementary, which is the whole point of combining them.
Why can't a bigger LLM fix hallucination?
Because the limit is mathematical rather than a matter of training budget. Xu, Jain and Kankanhalli show that LLMs cannot learn all computable functions and will therefore hallucinate when used as general problem solvers, and a 2025 survey of the fundamental limits of scaling describes hallucination as an intrinsic computational and statistical limit that no present architecture or scale can remove. Scaling raises average quality. It does not deliver a guarantee, and enterprise decisions need one.
How does a knowledge graph reduce LLM hallucination?
A knowledge graph stores facts as entities and typed relationships, each with a source and a date, so the system looks facts up instead of asking the model to remember them. The LLM's job shrinks to understanding the question and phrasing the answer; the facts come from the graph, and every claim can carry a citation. Many hallucination problems are grounding problems in disguise: the model was asked a factual question with no reliable data underneath it.
Is RAG the same as neuro-symbolic AI?
Retrieval-augmented generation is one step in that direction, not the whole thing. RAG retrieves relevant documents and hands them to the model, which grounds the answer in real text but still leaves the reasoning to the LLM. A neuro-symbolic system also moves the decision itself into a deterministic layer, such as a rules engine or solver, so the logic is applied the same way every time and the rule that fired can be named. Grounding improves the facts; the symbolic layer guarantees the decision.
When is an LLM on its own good enough?
When a wrong answer is cheap and a person reviews the output before anything irreversible happens: drafting, summarizing, search, support triage, brainstorming and first-pass classification. Those are tasks where language is the whole job, and an LLM is the best tool anyone has ever had for language. The moment the output becomes a decision that costs money, affects safety or has to survive an audit, route the decision to something deterministic and keep the LLM as the interface.
How does neuro-symbolic AI help with AI audits and ISO 42001?
Auditors and enterprise AI committees ask how a decision was made, which data it used, and whether it can be reproduced. A neuro-symbolic system answers all three by design: the facts came from a named source, the rule that fired has a version, and the numeric result is the same every time you run it. That is the evidence an AI management system under ISO 42001 exists to produce, and it is far easier to generate from a structured decision layer than from a model's free-text explanation.
Is neuro-symbolic AI new?
No. Symbolic AI was the dominant approach from the 1960s through the 1980s, in the form of expert systems built from hand-written rules. DARPA calls that era the first wave of AI, statistical learning the second, and contextual reasoning the third wave it now funds. What is new is that the neural side matured enough to carry its share: models can finally read messy human input well enough to sit in front of the symbolic systems that have quietly made credit, fraud and scheduling decisions for decades.
Written by the team behind The Security Podcast of Silicon Valley

Put it into practice.