Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Beyond the Model: Securing What Your Agent Actu...

Beyond the Model: Securing What Your Agent Actually Depends On [AAIF Bengaluru]

As AI adoption accelerates, so does a largely underestimated attack surface — the AI supply chain. Unlike traditional software, an AI system's supply chain spans training datasets, pre-trained models, ML frameworks, third-party packages, and inference pipelines. Each layer is a potential entry point for adversaries.
This talk maps the full AI supply chain and examines real-world attack vectors: data poisoning, backdoored models distributed via public hubs, malicious ML packages, unsafe model deserialization, and prompt injection through compromised data sources. We'll look at documented incidents and emerging research that show these are not theoretical risks.
The highlight is a live demo of OpenSSF Model Signing (OMS) — an industry standard developed by the OpenSSF AI/ML Working Group, backed by Google, NVIDIA, and HiddenLayer. The demo will show how to sign a model at training time and verify its integrity before deployment — a critical step when the team training a model is rarely the same one deploying it.

Avatar for Shivam Saraswat

Shivam Saraswat

September 25, 2026

More Decks by Shivam Saraswat

Other Decks in Technology

Transcript

  1. AAIF BENGALURU Beyond the Model Securing What Your Agent Actually

    Depends On LIVE DEMO — OPENSSF MODEL SIGNING (OMS) Shivam Saraswat sha256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08
  2. You just downloaded a model with 50,000 downloads. How do

    you know it's what the author actually trained?
  3. 00 · INTRO Shivam Saraswat Senior Product Security Engineer, PayPal

    4.5+ years leading DevSecOps and cybersecurity at PayPal, IKEA, Tekion, and BreachLock. Container Security Supply Chain Security CI/CD Security @shivamsaraswat
  4. 00 · INTRO What You'll Walk Away With 01 Why

    an agent's dependency chain is wider than “just the model" 02 A quick tour of what actually goes wrong when you trust it blindly 03 A working defense you can adopt — model signing 04 A live demo of cryptographic verification, start to finish
  5. 00 · INTRO Why This Matters Now You're not shipping

    one model anymore — you're shipping an agent that dynamically calls models, tools, and MCP servers Every one of those is something you didn't build and can't fully inspect We instrument agents for reliability and evals. We rarely ask: can I trust the pieces underneath? Trust, right now, is mostly assumed — not verified
  6. 01 · YOUR AGENT'S DEPENDENCIES A Model's Supply Chain S1

    Data S2 → Training S3 → Model artifacts S4 → Distribution — hubs ▲ EVERY STAGE IS AN ENTRY POINT S5 → Deployment S6 → Inference
  7. 01 · YOUR AGENT'S DEPENDENCIES The Agent's Supply Chain Is

    Wider Than You Think A MODEL'S data → training → artifacts → distribution → deployment → inference AN AGENT'S — ALL OF THAT, PLUS + MCP + TOOLS + FRAMEWORK Every MCP server it's allowed to call Every tool or plugin it invokes The orchestration framework gluing it together
  8. The team that builds a model is almost never the

    team that runs it. That gap — between creation and consumption — is the attack surface.
  9. 01 · YOUR AGENT'S DEPENDENCIES Not Like the Dependencies You

    Know A PACKAGE YOU INSTALL A MODEL YOU PULL Source code you can read and review Opaque weights — you can't code-review 7B parameters Established SBOM and scanning tooling Mature signing practices Unsafe formats that execute code on load Signing and provenance still emerging
  10. 02 · WHAT GOES WRONG What Actually Goes Wrong Four

    things that bite you when you pull an open-source model into production 01 02 03 04 Loading it can run code It can look clean and still be backdoored It can be a fake with a familiar name Malicious model serialization Model Backdoors (Trojan Models) Typosquatting Its training data can be the problem before the model even exists Data poisoning MORE DETAIL → blog.shivamsaraswat.com/ai-static-threats
  11. 01 / 04 Loading It Can Run Code You know

    how an npm install with a postinstall script can run arbitrary code before you've written a line of yours? Some model formats work the same way — loading the weights can execute a payload, not just read numbers. Legacy formats (pickle-based .pt / .pkl) are the usual culprit Safer alternatives exist (e.g. safetensors) — but plenty of old files are still out there PyTorch has a weights_only=True parameter
  12. 02 / 04 It Can Look Clean and Still Be

    Backdoored A model can pass every benchmark you throw at it — and still misbehave on a specific trigger nobody tested for. Hidden behavior activates only under narrow conditions Standard eval suites won't catch it — passing your benchmarks isn't the same as being trustworthy
  13. 03 / 04 It Can Be a Fake With a

    Familiar Name Same game as typosquatted npm/PyPI packages — now aimed at model hubs. ✓ meta-llama / Llama-3-8B ✗ rneta-llama / Llama-3-8B One character off from a name you trust Copy-paste a model ID wrong once, and you've pulled someone else's model into your agent
  14. 04 / 04 The Problem Can Start Before the Model

    Exists Poisoned training data → a model that's compromised from the start, regardless of how carefully you deploy it Hard to fully audit web-scale datasets
  15. 03 · THE DEFENSE Enter: Model Signing Sign the model

    when it's trained. Verify it every time it's used. INTEGRITY ORIGIN PRECEDENT The model hasn't been tampered with It came from who it claims to Standard practice for software — now for AI
  16. 03 · THE DEFENSE OpenSSF Model Signing (OMS) An industry

    standard from the OpenSSF AI/ML Working Group BACKED BY INTERFACE Google · NVIDIA · HiddenLayer Library + CLI — any model format, any size FLEXIBLE PKI DETACHED SIGNATURES Sigstore, self-signed certs, or key pairs Sigstore Bundle Format — the model is never modified ADOPTION Already integrated into major model hubs
  17. 03 · THE DEFENSE How OMS Works 01 · SIGN

    02 · SIGNATURE BUNDLE 03 · VERIFY At training time Detached — travels beside the model At deployment Hash every file, sign the manifest → Manifest of file hashes · optional metadata (version, data source, hardware) · signature over an in-toto statement → Every time, before the model is used
  18. 03 · THE DEFENSE Live Demo OpenSSF Model Signing in

    action $ sign the model $ verify it $ tamper with one byte $ verify again ✓ VERIFIED ✗ FAILED
  19. 03 · THE DEFENSE What Just Happened 01 One changed

    byte → verification breaks 02 No signature = no trust 03 This runs in CI/CD — sign at training time, verify at deploy time 04 Non-intrusive — the model file itself is untouched
  20. 04 · IN PRACTICE A Secure Model Pipeline in Practice

    INBOUND TRANSFORM 1 Pull an open-source model 4 Sign — your internal approved OUTBOUND 7 Sign the new model weights attestation 2 Verify the publisher's signature 8 Publish to your internal registry 5 Train or fine-tune — do your work 3 Scan for static threats 9 Consumers verify the signature 6 Scan again — the output is a new artifact SIGN VERIFY before use
  21. 04 · IN PRACTICE Signing Is One Layer INBOUND →

    TRANSFORM → OUTBOUND — defense in depth for an agent's dependency chain Model scanning — catch unsafe serialization Provenance and lineage tracking Authenticated, verified MCP servers and tools — not just the model Secure pipelines — least privilege, artifact integrity + Model signing — today's focus
  22. 04 · IN PRACTICE Start Monday Three things you can

    do coming week 01 02 03 Before pulling a model or MCP server into your agent, check whether it's signed Pilot OMS signing on one model in your own pipeline If you're shipping a fine-tuned model or a custom MCP server — sign it. You're the upstream now.
  23. Trust in AI must be verifiable — not assumed. Thank

    you. Questions? sha256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08 GET THE SLIDES
  24. Let's Connect WEB X / TWITTER shivamsaraswat.com @thecybersapien LINKEDIN GITHUB

    shivamsaraswat shivamsaraswat sha256: 9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08