About
Richard Feng — web scraping engineer. I turn protected websites into clean, RAG-ready structured data for AI agents, retrieval pipelines, and LLMs.
Who I am
I’m Richard Feng, a freelance web-automation engineer. I specialize in web scraping, data extraction, and API reverse engineering — turning complex, protected websites into clean, structured data that AI systems can actually use.
My toolkit spans Node.js (TypeScript), Python, Go, and Java, with deep expertise in Crawlee, Playwright, and Cheerio.
What I do
I build and maintain the production-grade Apify actors in this catalog. Every actor outputs clean, RAG-ready JSON — schema-consistent, deduplicated, and ready to drop into a vector store, retrieval pipeline, or training set. Most public-data actors need no API key. My work spans five areas:
E-Commerce & retail data
Scrapers for major beauty, fashion, and retail platforms including Sephora (storefronts in the US, Canada, the EU, the Middle East, the UK, India, and Asia-Pacific), Ulta Beauty, Farfetch, Lululemon, Boohoo, Macy’s, and Talabat Mart, plus universal scrapers for any Shopify or WooCommerce store.
Public, financial & legal data
Official U.S. datasets as clean, RAG-ready JSON — SEC EDGAR filings & XBRL financials, USAspending.gov federal awards, CourtListener court opinions, and ClinicalTrials.gov studies.
Business, jobs & lead data
Indeed job postings and employer data, Shopify store leads, every Google ad a competitor runs, and virtual-mailbox (CMRA) locations.
Social & media intelligence
Bluesky posts, profiles, and interactions via the AT Protocol, plus a YouTube subtitle & transcript scraper that outputs JSON, SRT, VTT, or LLM-ready text in 100+ languages, and a TikTok transcript scraper.
Developer utilities
Tools beyond scraping — SEO & structured-data auditing and web-to-PDF/image rendering.
Specialties
- Reverse engineering private APIs — turning undocumented mobile/web endpoints into reliable data sources
- Anti-bot bypass — Cloudflare, Akamai WAF, and custom protections
- RAG-ready output — normalized, schema-consistent JSON built for retrieval and grounding
- Multi-region scraping — locales, currencies, and compliance handled
- High-reliability systems — extractors built to keep working unattended at scale
Tech stack
| Category | Technologies |
|---|---|
| Languages | TypeScript, Python, Go, Java |
| Scraping | Crawlee, Playwright, Cheerio, Parsel, got-scraping |
| Anti-bot | Fingerprint generators, TLS fingerprinting, session rotation |
| Infrastructure | Apify Platform, Docker, GitHub Actions |
| Testing | Vegeta, custom load-testing frameworks |
Work with me
I offer custom scraping solutions, RAG data-pipeline engineering, and ongoing data extraction services. If your AI needs the real web as clean data — let’s talk.