
Closed
Posted
I need a full-scale crawl of more than ten million Walmart product pages focused strictly on two elements: the complete product description (title, bullet points, long form copy, specifications) and every image available in the gallery for each SKU. Price, availability, reviews, or other metadata are not required at this stage, so the scraper can concentrate on content-rich fields and high-resolution media. The raw HTML is not necessary; I would like the text cleaned and structured (JSON or CSV) and the images downloaded or stored as direct links, organised by SKU. Because of the size of the catalogue, I expect the solution to run in parallel—cloud-based orchestration with Scrapy, Python requests, Playwright or a comparable headless browser is fine as long as it can bypass typical anti-bot measures and stay within Walmart’s published usage policies. Deliverables: • A repeatable scraper/spider with clear README. • The initial dataset covering 10M+ SKUs: one file with structured descriptions, a mirrored directory or bucket containing the images. • A brief performance report outlining crawl speed, error rate and any throttling safeguards. I will validate by spot-checking record count against random category URLs and confirming image paths resolve correctly.
Project ID: 40656230
56 proposals
Remote project
Active 4 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
56 freelancers are bidding on average $42 USD/hour for this job

Hi there, I understand you need to extract comprehensive product data from 10+ million Walmart pages focusing on descriptions and gallery images. This requires handling massive scale while respecting rate limits and avoiding detection. I'll build a robust distributed scraping system using Python with Scrapy for efficient crawling and concurrent requests. The architecture will include rotating proxies, intelligent delays, and retry mechanisms to handle Walmart's anti-bot measures. Each product page will be parsed to capture complete titles, bullet points, long-form descriptions, specifications, and all gallery images through multiple selectors to ensure comprehensive coverage. The system will feature real-time monitoring, error logging, and checkpoint recovery to handle interruptions during the large-scale operation. Data will be stored in structured format with proper deduplication and validation checks. For 10 million products, I'll implement horizontal scaling across multiple servers with load balancing to maintain steady progress while staying within acceptable request rates. Best Regards, Khorshed Alam, RS Software
$1,250 USD in 40 days
9.5
9.5

Dear , We carefully studied the description of your project and we can confirm that we understand your needs and are also interested in your project. Our team has the necessary resources to start your project as soon as possible and complete it in a very short time. We are 25 years in this business and our technical specialists have strong experience in PHP, Python, Data Processing, Web Scraping, Software Architecture, Scrapy, API, Data Collection and other technologies relevant to your project. Please, review our profile https://www.freelancer.com/u/tangramua where you can find detailed information about our company, our portfolio, and the client's recent reviews. Please contact us via Freelancer Chat to discuss your project in details. Best regards, Sales department Tangram Canada Inc.
$25 USD in 5 days
8.9
8.9

Hello, I your "Extract 10M Walmart Product Data" project description in detail and undertood your requirements. I've worked on many PHP projects in recent times. So I am confident on achieving your expected Goals. Please initiate a communication thread to discuss further and start with the project. ⭐ 5.0/5 from a recent client: "A more professional version: “Excellent work! The job was completed within the committed timeline. Great quality, professionalism, and timely delivery. Highly appreciated and recommended.”" Final timeline and cost will be confirmed in chat after a complete understanding and documentation of the project expectations in detail.
$15 USD in 1 day
7.6
7.6

Extracting over ten million product pages from Walmart's catalog presents unique challenges, particularly around compliance and data structure. I would develop a robust web scraper in Python, leveraging tools like Puppeteer or Selenium for headless browsing, ensuring adherence to Walmart’s usage policies while efficiently capturing high-quality content and images. I am skilled in web scraping, JSON/CSV data structuring, and handling large-scale data processing. My track record includes a 4.9-star rating across 200 client reviews and 220 projects completed, showcasing my commitment to delivering quality results. What specific performance metrics are you expecting in the final performance report?
$25 USD in 14 days
7.4
7.4

Hi there, I understand you need to build a scalable data pipeline to extract product descriptions and all associated gallery images for over 10 million Walmart SKUs. Operationally, this involves a distributed system that manages a vast URL queue, orchestrates parallel scraping workers to fetch and parse content, navigates anti-bot measures, and then routes cleaned text and image assets into separate, structured deliverables. Technical approach: A cloud-native solution using Python with Scrapy, orchestrated on AWS. Product URLs will be managed in an SQS queue, feeding a fleet of containerized scraping workers on ECS. We'll use a premium rotating residential proxy service and dynamic user-agent/header management to ensure high success rates. Core modules: - URL Discovery & Queueing: Populates and manages the master list of 10M+ target URLs. - Distributed Crawling Workers: A fleet of spiders that fetch pages, parse description fields, and extract all image source URLs. - Data Processing & Storage: A pipeline to clean and structure text into a JSONL file while asynchronously downloading images to an S3 bucket, organized by SKU. - Monitoring & Throttling: Manages crawl speed, error rates, and implements retry logic. Implementation strategy: We'll start with a pilot crawl on a single large category to fine-tune selectors and the anti-bot strategy. Once validated, we will scale the worker fleet to execute the full 10M+ SKU crawl. The process will be monitored continuously to manage performance and error rates. The final delivery will include the dataset, image store, and the documented scraper source code. Regards, Rohit
$17 USD in 45 days
8.0
8.0

★★★ TOP 1% IN FREE LANCER WORLD ★★★ ★★★ 20+ Year Experience in IBD being CMD★★★ ★★★ 200+ Country Satisfied Clientele ★★★ ★Linkedin★ ★Data Entry★ ★Business Plans★★★ ★★★ Operational Strategic planner Customer Support 24*7★★★ ★★★Excel/Word Operation★★★ ★★★Chat Support★★★ ★★★Calling Support★★★ ★★★Business Plans / Marketing Strategy ★★★ * Digital Marketing★★★ ★★★Social Media Marketing ★★★ ★★★Internet Marketing ★★★ ★★★Any type of Data Projects★★★ ★★★★★★ Regards, ★★★CMD★★★ ★★★PVSYS GROUP (INDIA)★★★ ★★★IF YOU THINK THEN I CAN★★★
$15 USD in 40 days
7.8
7.8

Hello Your need for a full-scale crawl of over ten million URLs presents a significant data challenge. I've successfully executed large-scale web scraping projects, delivering clean, structured data for clients needing millions of records. Let's discuss how I can efficiently extract and organize the information you require. Giáp Văn Hưng
$25 USD in 7 days
6.9
6.9

Hi, I am interested to work on this Walmart Data Extraction project.I have done similar work before so I assure you that I can do this job perfectly within required time and reasonable budget. Message me here & LET'S GET STARTED THE WORK. I am looking forward to an early and positive response. Regards, Shalu
$20 USD in 40 days
7.0
7.0

Hi, I build a large variety of marketplaces / e-sommerce website data extractors, and i can help you with your task. I can deliver a both the data and the tool. I'm not ok with hourly rate i'd like to work on project-fixed. feel free to message me for more details. Regards!
$20 USD in 40 days
6.0
6.0

I can help you build a reliable, repeatable pipeline that extracts 10M+ Walmart product descriptions and images without getting blocked. My approach focuses on fault tolerance and data integrity at scale. I'll architect a distributed Scrapy cluster that prioritizes content-rich fields (title, bullets, long-form copy, specs) and high-res media, strictly ignoring price and reviews to maximize throughput. The system will feature automatic retry logic with exponential backoff, rotating proxy management, and adaptive throttling to navigate anti-bot measures while respecting Walmart's terms. For the 10M SKU volume, I'll implement a checkpointing system so the crawl can pause and resume seamlessly without data loss. The output will be clean, normalized JSON/CSV organized by SKU, with images stored in a mirrored directory structure or cloud bucket for easy validation. I'll also include a performance report detailing crawl speed, error rates, and the specific throttling safeguards used to maintain stability. The final deliverable is a fully documented, production-ready scraper that you can run repeatedly, plus the initial dataset ready for your spot-checking.
$20 USD in 40 days
6.1
6.1

Hi, I am a Python web scraping developer with 8 years of experience in software development, with a strong background in large-scale data extraction and distributed crawling. I am familiar with Python, Scrapy, Requests, Playwright, multiprocessing, cloud workers, JSON, CSV, object storage, image downloading, retry queues, rate limiting, and large-volume data pipelines. I can build a repeatable crawler focused specifically on Walmart product descriptions, specifications, bullet content, and full image galleries, with SKU-based structured output and image organization. For 10M+ products, I would design the crawl around parallel workers, checkpointing, deduplication, retries, throttling, and detailed error logging so interrupted jobs can resume safely while respecting the available site/API access rules. I will also provide crawl-speed, failure-rate, and throttling metrics with the final pipeline. I'm an individual freelancer and can work in any time zone you prefer. Please contact me with the best time for you to have a quick chat. Looking forward to discussing more details. Thanks. Emile.
$20 USD in 40 days
5.8
5.8

Extracting and structuring data from 10 million Walmart SKUs is a big job but absolutely doable with the right approach. I’ve handled large-scale scrapes before where focusing only on essential product content and images was key, which helped keep the pipeline efficient and avoid overload. I’d build a cloud-based Scrapy spider that runs in parallel across categories to speed things up. To handle throttling and anti-bot measures, I’d implement rotating proxies and user-agent switching, and add polite delays where needed. Cleaning and structuring the extracted text into JSON or CSV will be part of the pipeline, along with downloading and organizing images by SKU in a storage bucket. For performance, I’d include logging for crawl rates, error counts, and response times to monitor efficiency and handle any slow-downs. Are there specific Walmart categories or SKU patterns you want prioritized initially? Also, would you prefer image downloads saved in AWS S3 or a similar service, or just direct links stored? Ready to start building this scalable scraper and deliver clean product info and images as you requested.
$15 USD in 7 days
5.9
5.9

Hi There, I have strong experience with Python web scraping, Scrapy, large-scale data extraction, parallel crawling, deduplication, and structured dataset generation. I can build a repeatable scraper that focuses specifically on the complete product description and every available gallery image for each Walmart SKU, cleaning and structuring the text into JSON or CSV and organising the image links or downloaded files by SKU. The crawler can run in parallel with appropriate throttling, retries, error handling, and anti-bot safeguards while respecting the site’s published usage policies. For a catalogue of this size, I would design the workflow for distributed execution rather than treating it as a single sequential crawl. I can also provide the README, initial 10M+ SKU dataset, organised image storage, and a performance report covering crawl speed, errors, and throttling. I’m ready to review the Walmart structure and confirm the most suitable crawling approach before starting the full run. Rishan Agnar Consulting
$15 USD in 40 days
5.8
5.8

Hello, I can build a Python scraper to crawl 10M+ Walmart SKUs, extracting titles, bullets, long descriptions, specifications, and full-resolution gallery images into structured JSON or CSV with images organised by SKU. I will run it with parallel workers and rotating proxies to handle anti-bot measures at scale. I can start today. For the run, I will add retry logic, throttling, and an error log feeding the performance report. Questions: 1) Do you have SKU or category seed lists, or should I discover them? 2) Store images in S3, GCS, or local? Looking forward to discussing further. Regards, Shayan.
$19 USD in 40 days
5.5
5.5

Hi, I can help you with this project. I have relevant experience with PHP and can handle the work from development to testing and delivery. I've reviewed your requirements and can provide a clean, reliable, and responsive solution. Let's discuss the details and get started. Best, Arslan Shahid
$15 USD in 7 days
5.8
5.8

Hi, I can build this at scale using Python/Scrapy with concurrent workers and structured JSON/CSV output. I’ll focus on: - Product descriptions and specifications - All available gallery image URLs - SKU-based organization - Retry/error handling and throttling - Deduplication and resumable crawling - Performance/error-rate reporting - Clean, reusable source code + README For 10M+ products, I’d recommend validating the architecture with a smaller batch first, then scaling horizontally once accuracy and throughput are confirmed. I can provide a sample run and performance benchmark before the full crawl.
$18 USD in 40 days
5.7
5.7

Hi Syaiful, I will build a parallel Scrapy/Playwright spider that crawls over 10 M Walmart SKUs, extracts titles, bullet points, long copy, specs, and all gallery images, delivering cleaned JSON/CSV and an image bucket plus README and performance report. I can deliver a functional prototype within 3 days. Ready to start now—shall I use your AWS bucket for image storage? Waiting for your response in chat! Best Regards.
$20 USD in 3 days
5.5
5.5

★•══•★ Hi client ★•══•★ I’m ready to build a robust scraper tailored to grab Walmart’s product descriptions and images cleanly and efficiently. I’ll focus on extracting all text fields and high-res images, organizing everything by SKU in neat JSON or CSV formats with image links or downloads. To tackle the scale of 10M+ SKUs, I’ll set up a cloud-based parallel crawling system using Scrapy combined with headless browsing where needed. This approach keeps things speedy while respecting Walmart’s usage policies and dodging anti-bot blocks. You’ll get a reusable spider with clear instructions, the full dataset neatly organized, plus a concise report on crawl speed, errors, and throttling methods to ensure smooth runs. I’ll double-check data integrity so you can trust the results. How would you like me to handle image storage—direct links or cloud bucket uploads? Best regards, Rico
$20 USD in 40 days
5.1
5.1

The hardest part of this job is reliably getting through ten million Walmart product pages without getting blocked, so I would use a distributed scraping setup. I’d orchestrate this with Scrapy, letting it handle the request queueing and error handling for each SKU. For image downloading, I’d use Python’s `requests` library to grab them directly after parsing the HTML with `BeautifulSoup` to find the gallery links, also parsing product descriptions into a structured JSON output. I would assume the product page structure is consistent enough that the CSS selectors I identify from an initial sample will hold true across the catalogue. My one question is about the rate of change for product data; will we need to re-crawl periodically to catch updates, or is this a one-time historical snapshot? I have 8 reviews on here, everything delivered on time and on the agreed price so far, plus Preferred Freelancer status. Once I have the answer to that question, I will send over a detailed plan for the infrastructure and estimated timeline.
$25 USD in 7 days
5.2
5.2

Nice to meet you ,The requirements of your project match my areas of work and skills, to introduce myself. My name is Anthony Muñoz and i am the lead engineer for DS Pro IT agency. I have worked for over 10 years as a Full-Stack and software development engineer and have successfully done multiple jobs. It will be a pleasure to work together to make your project. Feel free to discuss about the project with me, greetings.
$47 USD in 40 days
4.9
4.9

Indonesia
Member since Mar 22, 2026
min €36 EUR / hour
$10-30 USD
₹750-1250 INR / hour
₹12500-37500 INR
$2-8 USD / hour
$2-8 USD / hour
$15-25 USD / hour
₹1500-12500 INR
₹600-1500 INR
₹12500-37500 INR
$2-8 USD / hour
$10-30 USD
$250-750 CAD
₹1500-12500 INR
$20-30 SGD / hour
₹600-1500 INR
₹600-1500 INR
$300-450 USD
₹12500-37500 INR
₹750-1250 INR / hour