Skip to content

Full-Spectrum Security Scanner with Trivy

We take a close look at the Trivy scanners for vulnerabilities, misconfigurations, and secrets with Ubuntu-centric guidance on performance tuning, security configurations, and scalability across Linux distributions.

Ocean diver
Photo by Jakob Boman on Unsplash

Deep Dive

Trivy is an open source, all-in-one security scanner by Aqua Security that has rapidly become a go-to tool in the DevSecOps toolkit [1]. It earned this reputation by combining multiple security checks such as vulnerability scanning, configuration auditing, and secret detection into a single easy-to-use command-line interface.

With a single binary and minimal setup, Trivy can be run on virtually any Linux distribution and integrated into continuous integration and continuous delivery (CI/CD) pipelines or cloud environments without friction. Its popularity is evident from a vibrant community (tens of thousands of GitHub stars) [2] and its adoption as the default scanner in platforms such as Harbor (the Cloud Native Computing Foundation (CNCF) image registry).

In an era of rising software supply chain threats, Trivy's approach addresses the need for continuous security checks at every stage of development and deployment. I walk you through Trivy's capabilities, how it works under the hood, and how to configure and optimize it on Linux for performance, security, and scale.

Main Capabilities

One reason Trivy is so valued is its breadth of coverage. It can scan container images (both local images and those in registries), the filesystems of running systems or virtual machines (VMs), remote code repositories, infrastructure-as-code (IaC) manifests, and even live Kubernetes clusters. In practice, this means a single tool can examine everything from your base operating system packages and application libraries to Dockerfiles, Terraform templates, and Kubernetes YAMLs, searching for any signs of security issues.

Trivy's primary function is to detect common vulnerabilities and exposures (CVEs) in software components. When pointed at a container image or a filesystem, it identifies vulnerable operating system packages (e.g., in Debian, Alpine, Ubuntu, RHEL, etc.), as well as vulnerabilities in application dependencies (language-specific packages like Pip/Pipenv for Python, npm/Yarn for Node.js, Maven for Java, etc.).

Under the hood, Trivy aggregates data from multiple security advisories and feeds (including distribution security bulletins and the US National Vulnerability Database (NVD)) into its database, which it updates regularly for high accuracy. This approach helps reduce false positives and ensures new CVEs are caught promptly.

The scanner maps each package in your image or system to known CVE entries and reports issues along with severity levels (e.g., LOW, MEDIUM, HIGH, CRITICAL) and available fixed versions. By generating a software bill of materials (SBOM) for the target and cross-referencing it against vulnerabilities, Trivy presents a detailed picture of risk in your software stack. Notably, Trivy is fast – the first scan might take a handful of seconds to download the latest vulnerability database, but subsequent scans are typically near-instant thanks to caching.

Beyond finding vulnerable software, Trivy can detect security misconfigurations in infrastructure definitions and container configurations. With a simple command, you can scan directories of IaC files (Terraform scripts, CloudFormation templates, Kubernetes manifests, Dockerfiles, etc.), and Trivy will automatically detect the file types and check them against built-in security policies. For example, it can catch whether a Dockerfile is using an outdated base image tag (latest instead of a specific version) or if a Kubernetes deployment manifest is missing recommended security settings. In one case, scanning a Dockerfile revealed a medium-severity misconfiguration: The base image was set to alpine:latest without a fixed tag, which is a security risk because the image can change over time.

Trivy shines light on such issues with explanations (e.g., When using a `FROM' statement you should use a specific tag to avoid uncontrolled behavior when the image is updated) and provides references to remediation guidance (e.g., links to the Aqua vulnerability database for misconfigurations). These checks help teams enforce best practices, from cloud environment settings to container security. The tool uses Open Policy Agent (OPA) under the hood with a library of Rego policies for common misconfigurations, and it even allows writing custom checks to tailor to your organization's standards, if needed.

Trivy also includes a scanner for sensitive information such as API keys, credentials, tokens, and other secrets that might be accidentally embedded in code or config files. When enabled, the secrets scanner will comb through file content in repositories or filesystems looking for patterns that match things like AWS secret keys, database connection strings, or private keys. The secrets scanner uses a combination of regex patterns and entropy analysis to detect high-entropy strings that likely represent secrets. Along with vulnerability and configuration issues, exposed secrets are listed in Trivy's output, making it a one-stop tool for a wide range of security checks.

As a bonus, Trivy can also identify software licenses in dependencies and flag license compliance issues (which is useful for organizations concerned about copyleft or proprietary licenses in their open source components). Additionally, Trivy can generate SBOMs for artifacts in standard formats (CycloneDX or SPDX), resulting in an inventory of components. This component is increasingly important for supply chain security and compliance. Trivy's recent versions even allow it to scan existing SBOM files for vulnerabilities, meaning if you have an SBOM (from Trivy or another tool), you can feed it to trivy sbom to identify any CVEs within it.

Under the Hood

Understanding how Trivy performs these scans helps in configuring and optimizing it effectively. Trivy uses a collection of scanners specialized for different purposes (vulnerabilities, configuration issues, secrets), and it supports multiple targets or artifact types to scan. When you run Trivy, you specify a target type as a subcommand (e.g., trivy image, trivy fs, trivy config, trivy repo, or trivy k8s) along with the subject (an image name, a path, a repository URL, or a cluster context). Trivy auto-detects what needs to be scanned and applies the relevant scanners.

For vulnerability scanning, Trivy maintains an internal vulnerability database (Trivy DB). The first time you run it, Trivy downloads this database, which includes aggregated CVE data from many sources (distro feeds for Linux packages, language-specific vulnerability databases). This initial download is on the order of a few tens of megabytes and is what makes the first scan take a bit longer (several seconds). After that, the database is cached locally, and Trivy updates it only if it's out of date. This design frees users from manually managing vulnerability feeds.

Trivy is largely stateless and automatically keeps its data current, avoiding the heavy setup of maintaining a separate database server. The scanning itself involves unpacking the target artifact: For a container image, Trivy will either pull it from the registry (if not available locally) or use the local docker or containerd daemon image if present, then analyze the filesystem layers to identify packages and libraries.

Trivy identifies operating system packages by package manager metadata (dpkg, rpm, apk indexes) and language libraries with manifest files such as pom.xml, package-lock.json, Gemfile.lock, and so on. Identified components are matched against the vulnerability database, and any findings are reported with details such as CVE IDs, severity, and fix versions. The scanning is performed efficiently in memory, and results can be output in various formats (table, JSON, SARIF, CycloneDX SBOM, etc.) depending on your needs.

For secrets detection, Trivy's mechanism runs patterns and entropy checks across text content. It's optimized to reduce false positives by ignoring common false patterns and using context (for instance, it knows to look for an AWS_SECRET_ACCESS_KEY= label near a high-entropy string to confirm it's likely a key). It's not foolproof, but it's a valuable net for catching mistakes. Secrets scanning can be resource-intensive (reading lots of files), so by default it's off for image and filesystem scans.

An important internal concept is caching. Trivy caches several things to improve performance: the vulnerability database (as mentioned) and the results of recent scans. It will cache image layer analysis so that if you scan the same image tag again, it doesn't need to redo all computations (unless you force refresh). You can explicitly clear the cache, for example, if you want to ensure a full fresh scan (perhaps after updating the database or Trivy version). The cache directory is configurable and can be moved to a shared path if you have multiple users or CI agents on the same host. In advanced setups, Trivy supports a Redis cache backend, which is particularly useful in client-server mode.

Optimizing Trivy

Trivy caches the vulnerability database and scan results to accelerate subsequent runs. By default, the database will update periodically, but if you are running many scans in a short time (say, in a CI environment with many parallel jobs), you can preload or host the database internally. One approach is to run

trivy image --download-db-only

on a schedule to fetch the latest database and share the ~/.cache/trivy directory among jobs.

In air-gapped environments with no Internet access, you can download (Trivy provides an offline database archive) and point to the database file for Trivy to use, ensuring scanning doesn't stall waiting for updates. Also, consider using the Trivy server mode: A central Trivy server can keep the database updated and serve all scanners, so each build agent doesn't redundantly download the same data.

Trivy will by default traverse through directories or image filesystems looking for all relevant files. In some cases, you might want to exclude certain paths to speed up scanning or avoid noise. Skipping large, irrelevant directories can save time. Likewise, if some files are known always to trigger false positives, those can be skipped. By tuning what Trivy looks at, you optimize runtime and focus on genuine risks.

Another performance consideration is deciding which scanners to enable. Running all scanners (vuln, config, secret) provides the most coverage but also does more work. If, for example, you only care about vulnerability scanning in a certain pipeline, you can stick to the default vuln-only scan, which will be faster than enabling everything.

Conversely, in a security audit job, you might run the full suite. Trivy gives flexibility to run the right level of scanning for context. Additionally, you can reduce output clutter by using severity filters or ignoring unfixed/unpatchable issues so that reports remain actionable. Although these don't significantly change scan time, they improve signal-to-noise, which indirectly improves efficiency by focusing attention on what matters.

If you need to scan a very large number of artifacts (like a registry with thousands of images), you can parallelize Trivy because it is just a command-line interface tool – run multiple Trivy processes in parallel, each on different targets. The main bottleneck might be the CPU and network for pulling images. Ensure the system has enough of these resources (Trivy can benefit from multiple cores, especially when scanning large images or lots of files, because it can parallelize some operations). For container registry scanning at scale, you might script a loop that lists images and spawn concurrent Trivy scans. In client-server mode, a beefy server can help distribute the load as well. In Kubernetes, Trivy Operator automates this process by spinning up scan jobs for each image in the cluster in parallel, which is efficient.

Conclusion

Trivy's rise to prominence in the IT security world is well-deserved. It addresses a pressing need for comprehensive, continuous security scanning that keeps up with today's fast-paced development and deployment practices.

Add ADMIN IT Infrastructure & Operations on Google