Theo dõi giá đối thủ
Thiết lập theo dõi giá tự động trên các website đối thủ bằng FourA API.
Những gì bạn sẽ xây dựng
Một script Python có thể:
- Lấy trang sản phẩm từ danh sách URL đối thủ
- Trích xuất dữ liệu giá từ HTML
- Ghi kết quả vào file CSV
- Chạy theo lịch trình
Điều kiện tiên quyết
- Một FourA API key (lấy tại đây)
- Python 3.8+
- Các package
requestsvàbeautifulsoup4
pip install requests beautifulsoup4
Bước 1: Xác định mục tiêu của bạn
Tạo danh sách các URL sản phẩm cần theo dõi:
targets = [
{"name": "Competitor A - Widget", "url": "https://competitor-a.com/widget", "selector": ".price"},
{"name": "Competitor B - Widget", "url": "https://competitor-b.com/products/widget", "selector": "[data-price]"},
{"name": "Competitor C - Widget", "url": "https://competitor-c.com/item/123", "selector": ".product-price span"},
]
Bước 2: Lấy trang qua FourA
import requests
import time
SINGLE_URL = "https://eu.api.foura.ai/api/single/"
BROWSER_URL = "https://eu.api.foura.ai/api/browser/"
API_KEY = "YOUR_API_KEY"
HEADERS = {
"X-API-Key": API_KEY,
"Content-Type": "application/json"
}
# A wait doesn't clear these: they reset with the period (credits, traffic)
# or at midnight UTC (the daily Browser allowance).
STOP_LIMITS = {"plan_limit_credits", "plan_limit_bandwidth", "plan_limit_browser_daily"}
def fetch_page(url, use_browser=False, attempts=3):
if use_browser:
endpoint, field = BROWSER_URL, "body"
payload = {"url": url, "timeout_ms": 15000}
else:
endpoint, field = SINGLE_URL, "data"
payload = {"method": "GET", "url": url, "unblocker": True}
for _ in range(attempts):
resp = requests.post(endpoint, headers=HEADERS, json=payload)
if resp.status_code != 429:
return resp.json().get(field, "")
limit = resp.headers.get("X-FourA-Limit")
if limit in STOP_LIMITS:
raise RuntimeError(f"Plan limit reached: {limit}")
time.sleep(int(resp.headers.get("Retry-After", "5")))
raise RuntimeError(f"Still rate limited after {attempts} attempts: {url}")
Một 429 với plan_limit_credits, plan_limit_bandwidth hoặc plan_limit_browser_daily trong X-FourA-Limit sẽ không tự hết bằng cách đợi vài giây, vì vậy hàm sẽ dừng lại thay vì thử lại. Xem Rate Limits để biết từng giới hạn và thời điểm thiết lập lại.
Bước 3: Trích xuất giá
from bs4 import BeautifulSoup
import re
def extract_price(html, selector):
soup = BeautifulSoup(html, "html.parser")
element = soup.select_one(selector)
if not element:
return None
# Extract numeric price from text like "$49.99" or "49,99 EUR"
text = element.get_text(strip=True)
match = re.search(r'[\d,.]+', text)
return float(match.group().replace(',', '.')) if match else None
Bước 4: Chạy và ghi log kết quả
import csv
from datetime import datetime
def monitor_prices():
timestamp = datetime.now().isoformat()
results = []
for target in targets:
html = fetch_page(target["url"])
price = extract_price(html, target["selector"])
results.append({
"timestamp": timestamp,
"name": target["name"],
"url": target["url"],
"price": price
})
print(f"{target['name']}: {price}")
time.sleep(1) # Be polite
# Append to CSV
with open("prices.csv", "a", newline="") as f:
writer = csv.DictWriter(f, fieldnames=["timestamp", "name", "url", "price"])
if f.tell() == 0:
writer.writeheader()
writer.writerows(results)
if __name__ == "__main__":
monitor_prices()
Bước 5: Lập lịch chạy
Chạy script mỗi giờ với cron:
crontab -e
# Add this line:
0 * * * * cd /path/to/project && python3 monitor.py >> monitor.log 2>&1
Mẹo
- Bắt đầu với single endpoint, chuyển sang browser nếu các trang sử dụng JavaScript rendering
- Thêm xử lý lỗi: các trang web có thể thay đổi bố cục. Ghi log các trường hợp thất bại riêng biệt.
- Luôn cập nhật selector: khi đối thủ thiết kế lại giao diện, hãy cập nhật CSS selector
- Tôn trọng các trang web: giãn cách các request, tránh giờ cao điểm, tuân thủ robots.txt
Các bước tiếp theo
- Chọn endpoint phù hợp: Chọn phương pháp tối ưu nhất
- Xử lý lỗi: Xử lý các trường hợp thất bại một cách linh hoạt
- Cào dữ liệu website động: Xử lý các trang hiển thị bằng JavaScript