~/utility/schema-markup-scraper

Schema Markup Scraper & SEO Auditor

Extract JSON-LD, Microdata, RDFa, Open Graph, and Twitter Cards from any URL with a comprehensive SEO audit scoring system.

seo TypeScriptCrawlee Global
proooxy/schema-markup-scraper — spec
categoryutility / seo
languageTypeScript
stackTypeScript, Crawlee
marketsGlobal
outputclean, RAG-ready JSON

key features

Structured data extraction — JSON-LD, Microdata, and RDFa

Social meta tags — Open Graph, Twitter Cards, Dublin Core

SEO analysis with 0-100 scoring

Canonical URL and hreflang validation

Author extraction for EEAT signals

LocalBusiness detection with 80+ subtypes

Image alt text audit

Breadcrumb schema validation

Geo tags and NAP extraction

use cases

  • Technical SEO auditing at scale
  • Structured data validation for websites
  • Competitive SEO analysis — compare schema markup across competitors
  • EEAT signal assessment for content sites
  • Local SEO auditing for businesses
  • Pre-launch SEO checklist validation

input parameters

ParameterTypeRequiredDescription
startUrlsarrayrequiredURLs to analyze
proxyobjectoptionalProxy configuration
maxRequestsPerCrawlnumberoptionalLimit total URLs to audit
maxConcurrencynumberoptionalParallel requests
extractMetaTagsbooleanoptionalExtract meta tags (default: true)
extractSeoAnalysisbooleanoptionalRun SEO analysis (default: true)
computeSeoScorebooleanoptionalCalculate 0-100 SEO score (default: true)
extractGeoDatabooleanoptionalExtract geo tags and NAP data

Output Example

 1{
 2  "url": "https://developers.google.com/search/docs/appearance/structured-data/product",
 3  "title": "Intro to Product Structured Data on Google | Google Search Central",
 4  "linkedData": [
 5    {
 6      "@context": "https://schema.org",
 7      "@type": "BreadcrumbList",
 8      "itemListElement": [
 9        { "@type": "ListItem", "position": 1, "name": "Search Central", "item": "https://developers.google.com/search" }
10      ]
11    }
12  ],
13  "openGraph": {
14    "site_name": "Google for Developers",
15    "type": "website"
16  },
17  "twitterCard": {
18    "card": "summary_large_image",
19    "has_large_image": "true"
20  },
21  "seoAudit": {
22    "score": 95,
23    "issues": [
24      { "severity": "warning", "code": "TITLE_TOO_LONG", "message": "Title is 122 chars (recommended: max 60)" }
25    ]
26  },
27  "headings": {
28    "headings": [
29      { "level": 1, "text": "Introduction to Product structured data" },
30      { "level": 2, "text": "Deciding which markup to use" }
31    ],
32    "h1Count": 1,
33    "issues": []
34  }
35}

faq

What structured data formats are supported?
JSON-LD, Microdata, and RDFa. The scraper also extracts Open Graph, Twitter Cards, and Dublin Core metadata.
How is the SEO score calculated?
The 0-100 score evaluates title tags, meta descriptions, heading hierarchy, image alt text, canonical URLs, mobile viewport, structured data presence, and more.
Can I audit multiple pages at once?
Yes, provide multiple URLs in startUrls. The scraper processes them in parallel for fast bulk auditing.

related in ~/utility

Run Schema Markup Scraper & SEO Auditor, or get a custom build

Start extracting on Apify in minutes, or hire me to build a bespoke scraper and RAG pipeline for your exact source and schema.

run on Apify get custom data