~/utility/schema-markup-scraper
Schema Markup Scraper & SEO Auditor
Extract JSON-LD, Microdata, RDFa, Open Graph, and Twitter Cards from any URL with a comprehensive SEO audit scoring system.
seo
TypeScriptCrawlee
Global
categoryutility / seo
languageTypeScript
stackTypeScript, Crawlee
marketsGlobal
outputclean, RAG-ready JSON
key features
Structured data extraction — JSON-LD, Microdata, and RDFa
Social meta tags — Open Graph, Twitter Cards, Dublin Core
SEO analysis with 0-100 scoring
Canonical URL and hreflang validation
Author extraction for EEAT signals
LocalBusiness detection with 80+ subtypes
Image alt text audit
Breadcrumb schema validation
Geo tags and NAP extraction
use cases
- Technical SEO auditing at scale
- Structured data validation for websites
- Competitive SEO analysis — compare schema markup across competitors
- EEAT signal assessment for content sites
- Local SEO auditing for businesses
- Pre-launch SEO checklist validation
input parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
startUrls | array | required | URLs to analyze |
proxy | object | optional | Proxy configuration |
maxRequestsPerCrawl | number | optional | Limit total URLs to audit |
maxConcurrency | number | optional | Parallel requests |
extractMetaTags | boolean | optional | Extract meta tags (default: true) |
extractSeoAnalysis | boolean | optional | Run SEO analysis (default: true) |
computeSeoScore | boolean | optional | Calculate 0-100 SEO score (default: true) |
extractGeoData | boolean | optional | Extract geo tags and NAP data |
Output Example
1{
2 "url": "https://developers.google.com/search/docs/appearance/structured-data/product",
3 "title": "Intro to Product Structured Data on Google | Google Search Central",
4 "linkedData": [
5 {
6 "@context": "https://schema.org",
7 "@type": "BreadcrumbList",
8 "itemListElement": [
9 { "@type": "ListItem", "position": 1, "name": "Search Central", "item": "https://developers.google.com/search" }
10 ]
11 }
12 ],
13 "openGraph": {
14 "site_name": "Google for Developers",
15 "type": "website"
16 },
17 "twitterCard": {
18 "card": "summary_large_image",
19 "has_large_image": "true"
20 },
21 "seoAudit": {
22 "score": 95,
23 "issues": [
24 { "severity": "warning", "code": "TITLE_TOO_LONG", "message": "Title is 122 chars (recommended: max 60)" }
25 ]
26 },
27 "headings": {
28 "headings": [
29 { "level": 1, "text": "Introduction to Product structured data" },
30 { "level": 2, "text": "Deciding which markup to use" }
31 ],
32 "h1Count": 1,
33 "issues": []
34 }
35}
faq
What structured data formats are supported?
JSON-LD, Microdata, and RDFa. The scraper also extracts Open Graph, Twitter Cards, and Dublin Core metadata.
How is the SEO score calculated?
The 0-100 score evaluates title tags, meta descriptions, heading hierarchy, image alt text, canonical URLs, mobile viewport, structured data presence, and more.
Can I audit multiple pages at once?
Yes, provide multiple URLs in startUrls. The scraper processes them in parallel for fast bulk auditing.
related in ~/utility
media
Universal Web Printer — URL & HTML to PDF, Image
Convert URLs and HTML to PDF, PNG, JPEG, or WebP.
open