inference-builder
Generate deployable Vision AI pipelines with high-performance video and streaming capabilities on NVIDIA GPUs. Create GPU-accelerated inference microservices using DeepStream, Triton, vLLM, TensorRT-LLM, or PyTorch backends.
By NVIDIA73 GitHub starsUpdated 2 months ago
Directory evidence
- Runtimes
- Claude Code
- Parsed components
- 1 skill or MCP entry
- Source updated
- Jul 14, 2026
- Manifest status
- Canonical path parsed
The directory validates manifest shape and source location. It does not execute the plugin or provide a security endorsement. Review the indexing methodology →
Install inference-builder for Claude Code
claude plugin marketplace add IchenDEV/agent-plugin-mkt
claude plugin marketplace update agent-plugin-marketplace
claude plugin install inference-builder@agent-plugin-marketplacePaste and run these commands in a terminal with Claude Code. They add and refresh the PluginsMP catalog, then install this plugin.
The installer fetches third-party code from the source repository shown on this page. This directory validates manifest structure and source location, but does not perform a security audit; review the manifest, components, and source before installing.
Get the source manually
git clone https://github.com/NVIDIA-AI-IOT/inference_builderClone the source repository, then follow its setup instructions to add the plugin to a compatible client. The repository root is the plugin root.
Plugin files
├── .claude-plugin/plugin.json└── skills/inference-builder/SKILL.md
Included Skills1
Generate deployable Vision AI pipelines with high-performance video and streaming capabilities on NVIDIA GPUs. Activate when users want to: create GPU-accelerated inference microservices or standalone apps for vision, video, or streaming workloads; write or edit pipeline YAML configs; build Docker images for GPU inference; work with models from NGC or HuggingFace; or deploy with DeepStream, Triton, vLLM, TensorRT-LLM, or PyTorch backends.
Plugin manifests1
{
"name": "inference-builder",
"description": "Generate deployable Vision AI pipelines with high-performance video and streaming capabilities on NVIDIA GPUs. Create GPU-accelerated inference microservices using DeepStream, Triton, vLLM, TensorRT-LLM, or PyTorch backends.",
"author": {
"name": "NVIDIA"
},
"keywords": [
"inference",
"nvidia",
"deepstream",
"triton",
"vllm",
"tensorrt-llm",
"pytorch",
"docker",
"gpu"
]
}For maintainers
If you maintain this plugin, link to this source-backed listing from your README so users can review its manifest and indexed components.
[inference-builder on Agent Plugins Marketplace](https://pluginsmp.com/plugins/inference-builder)