Agent Plugins Marketplace
All plugins

inference-builder

Generate deployable Vision AI pipelines with high-performance video and streaming capabilities on NVIDIA GPUs. Create GPU-accelerated inference microservices using DeepStream, Triton, vLLM, TensorRT-LLM, or PyTorch backends.

Claude Code1 Skill

By NVIDIA73 GitHub starsUpdated 2 months ago

Directory evidence

Runtimes
Claude Code
Parsed components
1 skill or MCP entry
Source updated
Jul 14, 2026
Manifest status
Canonical path parsed

The directory validates manifest shape and source location. It does not execute the plugin or provide a security endorsement. Review the indexing methodology

Install inference-builder for Claude Code

Installs for the current user
claude plugin marketplace add IchenDEV/agent-plugin-mkt
claude plugin marketplace update agent-plugin-marketplace
claude plugin install inference-builder@agent-plugin-marketplace

Paste and run these commands in a terminal with Claude Code. They add and refresh the PluginsMP catalog, then install this plugin.

The installer fetches third-party code from the source repository shown on this page. This directory validates manifest structure and source location, but does not perform a security audit; review the manifest, components, and source before installing.

Get the source manually
git clone https://github.com/NVIDIA-AI-IOT/inference_builder

Clone the source repository, then follow its setup instructions to add the plugin to a compatible client. The repository root is the plugin root.

Plugin files

inference-builder/
├── .claude-plugin/plugin.json
└── skills/inference-builder/SKILL.md

Included Skills1

inference-builderskills/inference-builder/SKILL.md

Generate deployable Vision AI pipelines with high-performance video and streaming capabilities on NVIDIA GPUs. Activate when users want to: create GPU-accelerated inference microservices or standalone apps for vision, video, or streaming workloads; write or edit pipeline YAML configs; build Docker images for GPU inference; work with models from NGC or HuggingFace; or deploy with DeepStream, Triton, vLLM, TensorRT-LLM, or PyTorch backends.

Plugin manifests1

.claude-plugin/plugin.json
{
  "name": "inference-builder",
  "description": "Generate deployable Vision AI pipelines with high-performance video and streaming capabilities on NVIDIA GPUs. Create GPU-accelerated inference microservices using DeepStream, Triton, vLLM, TensorRT-LLM, or PyTorch backends.",
  "author": {
    "name": "NVIDIA"
  },
  "keywords": [
    "inference",
    "nvidia",
    "deepstream",
    "triton",
    "vllm",
    "tensorrt-llm",
    "pytorch",
    "docker",
    "gpu"
  ]
}

If you maintain this plugin, link to this source-backed listing from your README so users can review its manifest and indexed components.

[inference-builder on Agent Plugins Marketplace](https://pluginsmp.com/plugins/inference-builder)