경쟁사 가격 모니터링

FourA API를 사용하여 경쟁사 웹사이트 전반에서 자동화된 가격 모니터링을 설정합니다.

구축할 내용

다음을 수행하는 Python 스크립트:

  1. 경쟁사 URL 목록에서 제품 페이지 가져오기
  2. HTML에서 가격 데이터 추출
  3. 결과를 CSV 파일에 기록
  4. 정해진 일정에 따라 실행

필수 조건

  • FourA API 키 (여기서 발급)
  • Python 3.8+
  • requests 및 beautifulsoup4 패키지
pip install requests beautifulsoup4

1단계: 대상 정의

모니터링할 제품 URL 목록을 생성합니다.

targets = [
    {"name": "Competitor A - Widget", "url": "https://competitor-a.com/widget", "selector": ".price"},
    {"name": "Competitor B - Widget", "url": "https://competitor-b.com/products/widget", "selector": "[data-price]"},
    {"name": "Competitor C - Widget", "url": "https://competitor-c.com/item/123", "selector": ".product-price span"},
]

2단계: FourA를 통해 페이지 가져오기

import requests
import time

SINGLE_URL = "https://eu.api.foura.ai/api/single/"
BROWSER_URL = "https://eu.api.foura.ai/api/browser/"
API_KEY = "YOUR_API_KEY"

HEADERS = {
    "X-API-Key": API_KEY,
    "Content-Type": "application/json"
}

# A wait doesn't clear these: they reset with the period (credits, traffic)
# or at midnight UTC (the daily Browser allowance).
STOP_LIMITS = {"plan_limit_credits", "plan_limit_bandwidth", "plan_limit_browser_daily"}

def fetch_page(url, use_browser=False, attempts=3):
    if use_browser:
        endpoint, field = BROWSER_URL, "body"
        payload = {"url": url, "timeout_ms": 15000}
    else:
        endpoint, field = SINGLE_URL, "data"
        payload = {"method": "GET", "url": url, "unblocker": True}
    for _ in range(attempts):
        resp = requests.post(endpoint, headers=HEADERS, json=payload)
        if resp.status_code != 429:
            return resp.json().get(field, "")
        limit = resp.headers.get("X-FourA-Limit")
        if limit in STOP_LIMITS:
            raise RuntimeError(f"Plan limit reached: {limit}")
        time.sleep(int(resp.headers.get("Retry-After", "5")))
    raise RuntimeError(f"Still rate limited after {attempts} attempts: {url}")

X-FourA-Limit에 plan_limit_credits, plan_limit_bandwidth 또는 plan_limit_browser_daily이 포함된 429은 몇 초 기다려도 해결되지 않으므로 함수가 재시도하는 대신 중지됩니다. 각 한도 및 재설정 시점은 Rate Limits를 참고하십시오.

Step 3: Extract Prices

from bs4 import BeautifulSoup
import re

def extract_price(html, selector):
    soup = BeautifulSoup(html, "html.parser")
    element = soup.select_one(selector)
    if not element:
        return None
    # Extract numeric price from text like "$49.99" or "49,99 EUR"
    text = element.get_text(strip=True)
    match = re.search(r'[\d,.]+', text)
    return float(match.group().replace(',', '.')) if match else None

4단계: 실행 및 결과 로깅

import csv
from datetime import datetime

def monitor_prices():
    timestamp = datetime.now().isoformat()
    results = []

    for target in targets:
        html = fetch_page(target["url"])
        price = extract_price(html, target["selector"])
        results.append({
            "timestamp": timestamp,
            "name": target["name"],
            "url": target["url"],
            "price": price
        })
        print(f"{target['name']}: {price}")
        time.sleep(1)  # Be polite

    # Append to CSV
    with open("prices.csv", "a", newline="") as f:
        writer = csv.DictWriter(f, fieldnames=["timestamp", "name", "url", "price"])
        if f.tell() == 0:
            writer.writeheader()
        writer.writerows(results)

if __name__ == "__main__":
    monitor_prices()

5단계: 일정 등록

cron을 사용하여 스크립트를 매시간 실행합니다.

crontab -e
# Add this line:
0 * * * * cd /path/to/project && python3 monitor.py >> monitor.log 2>&1

팁

  • 단일 endpoint로 시작하고, 페이지에서 JavaScript 렌더링을 사용하는 경우 브라우저로 전환하십시오.
  • 오류 처리 추가: 사이트 레이아웃은 변경될 수 있습니다. 실패 항목을 별도로 기록하십시오.
  • 선택자 최신 상태 유지: 경쟁 사이트가 개편되면 CSS 선택자를 업데이트하십시오.
  • 대상 사이트 보호: 요청 간격을 두고, 피크 시간대를 피하며, robots.txt를 준수하십시오.

다음 단계

최근 업데이트: 2026년 9월 23일