You detect a website's tech stack by reading what the site sends to every visitor. That means the HTML source, the HTTP response headers, the script and stylesheet URLs, the cookies, and the DNS records of the domain. Each one carries fingerprints. A path like /wp-content/ points to WordPress. A cf-ray header points to Cloudflare. A script loaded from cdn.shopify.com points to Shopify. You can check all of this for free with your browser's developer tools or a single curl command.
The catch is that you only see the outside of the system. Frontend frameworks, content management systems, analytics tags, CDNs and hosting are often visible. Databases, backend languages, internal tools and anything behind a login are mostly invisible. Every detection result is an inference from public signals, not a confirmed inventory. The rest of this article explains each signal, how to read it, and where it misleads you.
Signal 1: the HTML source
Open any page, right-click and choose "View page source". Much of the stack is visible here.
The generator meta tag. Many content management systems add a tag like <meta name="generator" content="WordPress 6.x">. The HTML standard defines generator as a standard metadata name that identifies the software used to produce the page. WordPress, Drupal, Joomla, Hugo and many other systems set it by default. Site owners can remove it, and many do.
Asset paths. File paths are harder to hide than meta tags. WordPress serves themes and plugins from /wp-content/ and core files from /wp-includes/. Next.js serves its bundles from /_next/static/. Shopify stores load assets from cdn.shopify.com. Plugin names often appear in the path itself, so a WordPress page can reveal its plugins through its stylesheet URLs alone.
Framework markers in the markup. Some frameworks leave marks in the rendered HTML. Angular adds an ng-version attribute to the root element of the app. Next.js sites built with the pages router include a script tag with the id __NEXT_DATA__. Nuxt sites often expose a window.__NUXT__ object. These markers are side effects of how the frameworks work, which makes them more reliable than a tag that exists only to announce the software.
Class names and comments. CSS class conventions and HTML comments left by plugins or page builders also give hints. They are weaker evidence, because class names get copied between projects and comments survive migrations.
Signal 2: HTTP response headers
Every response from a server includes headers that the browser normally hides. You can see them in the Network tab of your developer tools, or with this command:
curl -sI https://example.com
The most direct header is Server. According to MDN's documentation of the Server header, it describes the software that the origin server used to handle the request. Typical values are nginx, Apache or cloudflare. MDN also warns against overly detailed values, because they can reveal information that helps attackers. That warning explains why many servers send a short value or none at all.
Other useful headers:
X-Powered-Byoften names the application layer, such as PHP or Express. Express sends it by default, and its own security guidance recommends turning it off.X-Generatoris sent by Drupal and some other systems.cf-rayindicates the request passed through Cloudflare. Other CDNs and hosts add their own headers, often with a recognizable prefix.Set-Cookiereveals session cookie names.PHPSESSIDsuggests PHP.JSESSIONIDsuggests a Java servlet container.ASP.NET_SessionIdsuggests ASP.NET.laravel_sessionsuggests Laravel.
Headers are easy to change or strip. A missing header proves nothing. A present header is usually honest, but it describes the layer that answered you, which may be a proxy and not the application behind it.
Signal 3: scripts and third-party requests
The Network tab in your browser's developer tools lists every request a page makes. This is where marketing and analytics tools show up. Analytics platforms, tag managers, chat widgets, A/B testing tools, payment providers, consent banners and font services each load from their own recognizable domains. The Chrome DevTools network documentation explains how to inspect and filter these requests.
JavaScript globals are a second layer. Many libraries register a global object on the page, which you can check in the console. jQuery exposes jQuery, and many analytics tools expose their own queue or data layer objects. Browser-based detectors rely heavily on this, because a global object proves the library actually ran. A script URL only proves it was requested.
Signal 4: DNS records
DNS tells you about infrastructure that the page itself does not show. You can query it for free with dig or nslookup.
- MX records show which provider handles the domain's email.
- TXT records often contain SPF entries and domain verification strings. These name email senders and SaaS vendors that the company has authorized.
- NS and CNAME records show the DNS provider and often the hosting platform or CDN.
DNS signals cover the company, not only the website. An SPF record may list a marketing email service that never appears in any page source.
Signal 5: well-known paths
Some systems expose predictable URLs. WordPress ships with a REST API, and the WordPress REST API handbook documents how it is exposed under /wp-json/. A valid JSON response there is strong evidence of WordPress, even when the generator tag is gone. Files like robots.txt and sitemap.xml also leak structure, because they often list platform-specific paths or name the plugin that generated them.
Only request paths that a normal visitor or crawler would request. Probing for admin panels or guessing private URLs goes past detection and into territory that a site's terms may prohibit.
Where detection goes wrong
No method here is complete. These are the common failure modes.
The backend is invisible. A visitor who is not logged in sees the delivery layer. You cannot see the database, the queue system, the internal services, the data warehouse or the CRM. A static frontend may sit on top of anything. If someone claims to know a company's database from the outside, treat it as a guess unless another source supports it.
Logged-in areas use different technology. The marketing site and the product are often separate systems. A company can run its public site on a hosted site builder and its application on a custom stack. Scanning the homepage tells you about the homepage.
Consent banners hide marketing tags. On many sites, analytics and advertising scripts load only after the visitor accepts cookies. An automated scan that does not click "accept" will miss those tools entirely. The result looks like a site with no analytics, which is rarely true.
Tag managers and server-side tagging hide vendors. A tag manager loads other tools dynamically, sometimes only on certain pages or events. With server-side tagging, the browser talks to the company's own domain, and the vendor never appears in the network log.
Bundling removes fingerprints. Modern build tools combine libraries into files with hashed names. A library compiled into a bundle has no recognizable URL and may expose no global object. Detection of small libraries on bundled sites is weak.
Proxies mask the origin. When a CDN sits in front of a site, the Server header and IP address belong to the CDN. The real web server and host stay hidden.
Leftovers cause false positives. Script tags outlive contracts. A snippet from a tool the company cancelled long ago may still sit in the template. A blog post that mentions a technology can also trigger naive pattern matching. A detected tool is not proof of a paying customer.
Bot protection changes the response. Some sites serve a challenge page to automated clients. A scanner that does not notice will report the technology of the challenge page, not the site.
Version numbers are unreliable. Versions come from generator tags and file query strings, which are often removed, cached or deliberately changed.
Free ways to get the same data
You do not need a paid tool for a small number of sites.
- Browser developer tools cover the HTML, headers, cookies, network requests and JavaScript globals. For one site you care about, this is the most accurate method, because you can accept the consent banner, click through pages and see what really loads.
curlanddigcover headers and DNS from the command line and are easy to script.- Open source fingerprint libraries and browser extensions automate the pattern matching described above. They work well for checking sites one at a time.
- The HTTP Archive publishes a free public dataset of crawled pages, including detected technologies, that you can query in BigQuery. See httparchive.org for the methodology. It is the better choice for market-level questions, such as how widely a CMS is used. It crawls on a fixed schedule and does not cover every site, so it is less suited to a custom list of domains.
- Asking directly. There is no official API that returns a website's tech stack, because no site is required to publish it. Engineering blogs, job listings and public documentation often name the backend technologies that no scan can see. For the backend, those sources beat any detector.
Manual checking is the better choice when you have a handful of sites, need high confidence, or care about the backend. Automated detection is the better choice when you have a long list of domains and need consistent, structured output, and you accept that each result is a set of public signals with the limits described above.
If you need this for hundreds or thousands of domains, our Website Tech Stack Detector reads the public HTML, headers and scripts of each site and returns the detected technologies as structured data, and you pay per result.
Want the signal instead of the raw filings? Get a free report preview. Prefer the tool to the write-up? Browse all data feeds or connect the free MCP server.