Tuesday, 29 September 2026 Login

Code Without Boundaries

BREAKING
API Ecosystem

AI security patch issues go undisclosed

AI security patch issues go undisclosed - ai security patch issues
Chris Goettl trained a Claude skill on patch data for months before discovering the system frequently invented details.

Enterprise security teams are trusting AI to rank Patch Tuesday vulnerabilities, yet vendors rarely disclose if the guidance is entirely human-generated or partially machine-made. Chris Goettl, Ivanti’s VP of product management for endpoint security, spent months training a Claude skill on the same patch data he had processed by hand for a decade. The system kept inventing details. “You see something blatantly wrong, and at some point, you just get fed up, and you say, did you just make that up,” Goettl told VentureBeat. “And it will literally tell you that it has no actual foundation or referenceable material. That was all me.”

Goettl’s skill applies Ivanti’s own threat-risk prioritization to raw patch data, runs on Anthropic’s platform, and draws only on published vendor advisories and Ivanti’s own spreadsheets. It does not touch customer data, and every output is a draft until a human approves it for release. The prioritized Patch Tuesday briefing reaches 500 to 700 webinar attendees each month. Goettl trained the system vendor by vendor, starting with Microsoft because its downloadable spreadsheets made that training the cleanest. Adobe was harder. Adobe does not provide a clean spreadsheet, so Goettl fed the skill individual page URLs and trained it to read Adobe’s site structure.

He configured the skill to flag deviations from each vendor’s release pattern, so known exploits, public disclosures, and abnormal CVE volumes all trigger the risk-tier escalation his customers depend on. Adobe’s release cadence came back wrong repeatedly. Office editions needed clarifying because Microsoft breaks Office into multiple edition families and the skill conflated them. The pattern persists even in production. The fabrication catches are why the review step exists. Goettl built the human gate around the failure modes he found during training.

Goettl and Todd Schell, Ivanti’s senior product manager for patch, had spent about 48 hours around every Patch Tuesday researching vendors and assembling the prioritized briefing their customers receive. That effort consumed at least 5% of their combined working time each month. Schell replayed the skill on two earlier months of hand-built data. Ivanti says its reviewers found 98% alignment, with a few edge cases to fix.

VentureBeat did not independently audit the data set, scoring method, or error categories. The August run was the first production month. Goettl was on vacation. The spreadsheet that used to take about four hours each for two people ran in under 30 minutes. September 8, at 973 CVEs by Ivanti’s count, was the skill’s second production month.

Inside the advisory chain

Microsoft stopped listing its CVEs in its Security Update Guide starting with July’s Patch Tuesday. Every Patch Tuesday count is now a parse. CVEs resolved in Microsoft’s September 8, 2026 Patch Tuesday, as published by each tracker. Sources: Tenable, BleepingComputer, Zero Day Initiative, Ivanti, Rapid7, SecurityWeek and Senserva, September 8 and 9, 2026. Chart: VentureBeat. Seven trackers published totals for the same Tuesday.

Tenable counted 964 CVEs. Senserva, which counts CVEs against KB articles rather than advisories, reached 1,169. That is a 205-CVE gap for the same set of patches. Two of those vulnerabilities were zero-days already exploited in the wild before the fix shipped. The biggest Patch Tuesday before 2026, by Ivanti’s June post, was 175 CVEs in October 2025. September’s gap alone exceeds the old record.

What AI changes is the layer between the raw data and the risk tier a customer acts on. Not every gap requires AI to explain. Scope definitions and counting methods produce a wide spread on their own.

Accountability in a black box

An independent read Kayne McGladrey, senior IEEE member, independent vCISO, and author of “Cyber Risk is a Myth,” has no ties to Ivanti and answered in writing about the practice, not any one vendor. “The vendor owes their customers one page, in writing, before a single priority task lands in anyone’s queue,” McGladrey wrote in an email to VentureBeat. “That page should cover where the model sits in the pipeline, whether deterministic code parses the CVEs or the LLM does the counting, who reviewed the output, and what fraction of a 400-row list got verified against the source rather than skimmed.” Asked whether a replay plus a human read before release is sufficient, his answer was no. “A replay checks whether this month looks like last month. It tells you nothing about correctness. A human reading it might not do much better, because the failure that matters is the CVE that never appears, and nobody skimming a machine-generated list will notice something’s missing.”

Schell compared the skill’s output against months he and Goettl had built by hand, not against last month’s shape. That comparison does not address McGladrey’s second point. The risk tiers were not independently checked, and a row that never appears is still invisible to a reviewer reading the rows that did. VentureBeat Pulse’s July 2026 Agent Reliability and Evals tracker surveyed 108 enterprises across its research panel. 49% had shipped an agent that passed internal evals and then failed in front of a customer. 37% let an autonomous agent deploy changes with no human in the loop. 13% fully trust automated evaluation. Among the 53 enterprises burned by an eval failure, just 4% still trust it.

Disclosure and the pressure to adapt

Goettl’s reasoning is that the judgment is the product. “What really matters is my subject matter expertise,” he told VentureBeat. “Knowing what’s important, what’s not, having that informed ability to make judgment calls, to give recommendations, to guide the tools to help make this happen.” The pressure driving adoption is real. CISA’s June 2026 Binding Operational Directive 26-04 set a three-day window for the highest-risk vulnerabilities at federal civilian agencies.

Ivanti published a Patch Tuesday post on June 9 captioned “Graph generated using Claude (Anthropic) on June 9, 2026, based on author-designed prompts and dataset by Chris Goettl.” No regulation or industry standard required the disclosure, and that caption has sat on Ivanti’s site for three months. Computer Weekly, Cyber Magazine, Help Net Security, Krebs on Security, and Rapid7 have all run Goettl’s numbers this year. None reported that the guidance is partly skill-produced, what the skill got wrong during training, or the review chain between it and release.

McGladrey’s one-page standard gives a patch team the language to hold any vendor accountable. Any SOC leader receiving Patch Tuesday guidance can ask three questions today and get an answer before October 13: What in your Patch Tuesday guidance is AI-generated? What was it validated against, and does that validation catch a missing row? Who reviews it before release, and is that disclosed? October 13 is the next Patch Tuesday. The parses begin again, and the total will depend on what each vendor chooses to count.

Tags:

Leave a Reply

Your email address will not be published. Required fields are marked *