Theo dõi giá đối thủ

Thiết lập theo dõi giá tự động trên các website đối thủ bằng FourA API.

Những gì bạn sẽ xây dựng

Một script Python có thể:

  1. Lấy trang sản phẩm từ danh sách URL đối thủ
  2. Trích xuất dữ liệu giá từ HTML
  3. Ghi kết quả vào file CSV
  4. Chạy theo lịch trình

Điều kiện tiên quyết

  • Một FourA API key (lấy tại đây)
  • Python 3.8+
  • Các package requests và beautifulsoup4
pip install requests beautifulsoup4

Bước 1: Xác định mục tiêu của bạn

Tạo danh sách các URL sản phẩm cần theo dõi:

targets = [
    {"name": "Competitor A - Widget", "url": "https://competitor-a.com/widget", "selector": ".price"},
    {"name": "Competitor B - Widget", "url": "https://competitor-b.com/products/widget", "selector": "[data-price]"},
    {"name": "Competitor C - Widget", "url": "https://competitor-c.com/item/123", "selector": ".product-price span"},
]

Bước 2: Lấy trang qua FourA

import requests
import time

SINGLE_URL = "https://eu.api.foura.ai/api/single/"
BROWSER_URL = "https://eu.api.foura.ai/api/browser/"
API_KEY = "YOUR_API_KEY"

HEADERS = {
    "X-API-Key": API_KEY,
    "Content-Type": "application/json"
}

# A wait doesn't clear these: they reset with the period (credits, traffic)
# or at midnight UTC (the daily Browser allowance).
STOP_LIMITS = {"plan_limit_credits", "plan_limit_bandwidth", "plan_limit_browser_daily"}

def fetch_page(url, use_browser=False, attempts=3):
    if use_browser:
        endpoint, field = BROWSER_URL, "body"
        payload = {"url": url, "timeout_ms": 15000}
    else:
        endpoint, field = SINGLE_URL, "data"
        payload = {"method": "GET", "url": url, "unblocker": True}
    for _ in range(attempts):
        resp = requests.post(endpoint, headers=HEADERS, json=payload)
        if resp.status_code != 429:
            return resp.json().get(field, "")
        limit = resp.headers.get("X-FourA-Limit")
        if limit in STOP_LIMITS:
            raise RuntimeError(f"Plan limit reached: {limit}")
        time.sleep(int(resp.headers.get("Retry-After", "5")))
    raise RuntimeError(f"Still rate limited after {attempts} attempts: {url}")

Một 429 với plan_limit_credits, plan_limit_bandwidth hoặc plan_limit_browser_daily trong X-FourA-Limit sẽ không tự hết bằng cách đợi vài giây, vì vậy hàm sẽ dừng lại thay vì thử lại. Xem Rate Limits để biết từng giới hạn và thời điểm thiết lập lại.

Bước 3: Trích xuất giá

from bs4 import BeautifulSoup
import re

def extract_price(html, selector):
    soup = BeautifulSoup(html, "html.parser")
    element = soup.select_one(selector)
    if not element:
        return None
    # Extract numeric price from text like "$49.99" or "49,99 EUR"
    text = element.get_text(strip=True)
    match = re.search(r'[\d,.]+', text)
    return float(match.group().replace(',', '.')) if match else None

Bước 4: Chạy và ghi log kết quả

import csv
from datetime import datetime

def monitor_prices():
    timestamp = datetime.now().isoformat()
    results = []

    for target in targets:
        html = fetch_page(target["url"])
        price = extract_price(html, target["selector"])
        results.append({
            "timestamp": timestamp,
            "name": target["name"],
            "url": target["url"],
            "price": price
        })
        print(f"{target['name']}: {price}")
        time.sleep(1)  # Be polite

    # Append to CSV
    with open("prices.csv", "a", newline="") as f:
        writer = csv.DictWriter(f, fieldnames=["timestamp", "name", "url", "price"])
        if f.tell() == 0:
            writer.writeheader()
        writer.writerows(results)

if __name__ == "__main__":
    monitor_prices()

Bước 5: Lập lịch chạy

Chạy script mỗi giờ với cron:

crontab -e
# Add this line:
0 * * * * cd /path/to/project && python3 monitor.py >> monitor.log 2>&1

Mẹo

  • Bắt đầu với single endpoint, chuyển sang browser nếu các trang sử dụng JavaScript rendering
  • Thêm xử lý lỗi: các trang web có thể thay đổi bố cục. Ghi log các trường hợp thất bại riêng biệt.
  • Luôn cập nhật selector: khi đối thủ thiết kế lại giao diện, hãy cập nhật CSS selector
  • Tôn trọng các trang web: giãn cách các request, tránh giờ cao điểm, tuân thủ robots.txt

Các bước tiếp theo

Cập nhật: 23 tháng 9, 2026