~/about

About

Richard Feng — web scraping engineer with 12+ years of experience. I turn protected websites into clean, RAG-ready structured data for AI agents, retrieval pipelines, and LLMs.

Who I am

I’m Richard Feng, a freelance web-automation engineer with 12+ years of coding experience. I specialize in web scraping, data extraction, and API reverse engineering — turning complex, protected websites into clean, structured data that AI systems can actually use.

My toolkit spans Node.js (TypeScript), Python, Go, and Java, with deep expertise in Crawlee, Playwright, and Cheerio. I’ve built production systems that handle millions of requests at >99% success.

What I do

I build and maintain 16 production-grade Apify actors serving over 3,100 users at a consistent >99% run success rate. Every actor outputs clean, RAG-ready JSON — schema-consistent, deduplicated, and ready to drop into a vector store, retrieval pipeline, or training set. Most public-data actors need no API key. My work spans four areas:

E-Commerce & retail data

Scrapers for major beauty and fashion platforms including Sephora (21 global storefronts), Ulta Beauty, Farfetch, Lululemon, Boohoo, Macy’s, and a universal Shopify scraper that works with any Shopify-powered store.

Official U.S. datasets as clean, RAG-ready JSON — SEC EDGAR filings & XBRL financials, USAspending.gov federal awards, CourtListener court opinions, and ClinicalTrials.gov studies.

Social & media intelligence

Bluesky posts, profiles, and interactions via the AT Protocol, plus a YouTube subtitle & transcript scraper that outputs JSON, SRT, VTT, or LLM-ready text in 100+ languages.

Developer utilities

Tools beyond scraping — SEO & structured-data auditing, web-to-PDF/image rendering, and high-performance load testing.

Specialties

  • Reverse engineering private APIs — turning undocumented mobile/web endpoints into reliable data sources
  • Anti-bot bypass — Cloudflare, DataDome, Akamai WAF, and custom protections
  • RAG-ready output — normalized, schema-consistent JSON built for retrieval and grounding
  • Multi-region scraping — locales, currencies, and compliance handled
  • High-reliability systems — extractors that hold >99% success at scale

Tech stack

CategoryTechnologies
LanguagesTypeScript, Python, Go, Java
ScrapingCrawlee, Playwright, Cheerio, Parsel, got-scraping
Anti-botFingerprint generators, TLS fingerprinting, session rotation
InfrastructureApify Platform, Docker, GitHub Actions
TestingVegeta, custom load-testing frameworks

Work with me

I offer custom scraping solutions, RAG data-pipeline engineering, and ongoing data extraction services. If your AI needs the real web as clean data — let’s talk.