동적 웹사이트 스크래핑
동적 웹사이트는 초기 페이지 로드 후 JavaScript를 사용하여 콘텐츠를 로드합니다. 이 가이드에서는 FourA의 브라우저 endpoint를 사용하여 이러한 사이트에서 데이터를 수집하는 방법을 설명합니다.
문제점
JavaScript 비중이 높은 웹사이트에 표준 HTTP request를 보내면 HTML 셸만 수신되고 실제 콘텐츠는 수신되지 않습니다. 필요한 데이터(상품 목록, 가격, 검색 결과)는 브라우저에서 페이지가 렌더링된 후 JavaScript를 통해 로드됩니다.
이는 React, Vue, Angular, Next.js와 같은 최신 프레임워크에서 점점 더 흔해지고 있습니다.
해결 방법: 브라우저 Request
FourA의 브라우저 endpoint(POST /api/browser/)는 다음 작업을 수행하는 Chrome 브라우저 인스턴스에서 URL을 엽니다.
- 페이지 로드
- 모든 JavaScript 실행
- 콘텐츠가 렌더링될 때까지 대기
- 완전히 렌더링된 HTML 반환
1단계: 필요한 요소 식별
request를 보내기 전에 브라우저에서 대상 페이지를 방문하고 DevTools(F12)를 사용하여 콘텐츠가 로드되었는지 확인할 수 있는 텍스트나 요소를 찾습니다. 예시:
- JS 렌더링 후 표시되는 제품 이름
- 렌더링된 HTML 내
product-grid와 같은 CSS 클래스 - 데이터가 로드될 때만 표시되는 "results"와 같은 텍스트 문자열
2단계: 브라우저 Request 전송
curl -X POST https://eu.api.foura.ai/api/browser/ \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products",
"timeout_ms": 15000,
"checkText": "product-grid"
}'
checkText 옵션은 페이지 로드가 완료된 후 "product-grid"가 페이지에 있는지 확인하도록 FourA에 지시합니다. 없는 경우 요청은 checkText:product-grid not found 오류와 함께 실패로 반환되며, 실패한 요청에는 요금이 청구되지 않습니다. checkText(으)로 인해 FourA가 텍스트를 더 오래 기다리지는 않습니다.
3단계: HTML 파싱
응답의 body 필드에 완전히 렌더링된 HTML이 포함되어 있습니다. 선호하는 라이브러리로 파싱하세요.
Python (BeautifulSoup)
import requests
from bs4 import BeautifulSoup
resp = requests.post("https://eu.api.foura.ai/api/browser/", headers={
"X-API-Key": "YOUR_API_KEY",
"Content-Type": "application/json"
}, json={
"url": "https://example.com/products",
"timeout_ms": 15000,
"checkText": "product-grid"
})
html = resp.json()["body"]
soup = BeautifulSoup(html, "html.parser")
for product in soup.select(".product-card"):
name = product.select_one(".product-name").text.strip()
price = product.select_one(".product-price").text.strip()
print(f"{name}: {price}")
Node.js (cheerio)
import * as cheerio from 'cheerio';
const resp = await fetch('https://eu.api.foura.ai/api/browser/', {
method: 'POST',
headers: { 'X-API-Key': 'YOUR_API_KEY', 'Content-Type': 'application/json' },
body: JSON.stringify({
url: 'https://example.com/products',
timeout_ms: 15000,
checkText: 'product-grid'
})
});
const { body: html } = await resp.json();
const $ = cheerio.load(html);
$('.product-card').each((i, el) => {
console.log($(el).find('.product-name').text(), $(el).find('.product-price').text());
});
문제 해결
여전히 빈 콘텐츠가 반환되나요?
- 페이지가 실제로 JavaScript 렌더링을 사용하는지 확인합니다 ("소스 보기"와 DevTools 비교 확인).
timeout_ms값을 늘립니다. 일부 페이지는 로드 속도가 느립니다.- 페이지에 인증 또는 cookie가 필요한지 확인합니다 (
cookies파라미터 사용).
인증 페이지가 표시되나요?
- 단일/HTTP request의 경우, 자동 IP 교체를 위해 proxy endpoint (
POST /api/proxy/)로 전환합니다. - browser request를 proxy를 통해 라우팅하려면 browser endpoint의
proxy파라미터를 전달합니다. proxy endpoint는 browser request가 아닌 단일/HTTP request만 래핑합니다.
proxy 파라미터는 자체 proxy 주소가 아니라 이전 response가 반환한 불투명 proxy ID를 받습니다. 먼저 POST /api/proxy/ (또는 POST /api/auto/) 호출을 한 번 실행하고, 반환된 proxy 필드를 읽은 후 고정합니다.
# Step 1: let FourA find a working exit. The response carries "proxy": "A1B2C3".
curl -X POST "https://eu.api.foura.ai/api/proxy/" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"maxTries": 5, "request": {"method": "GET", "url": "https://example.com/products"}}'
# Step 2: render through that same exit.
curl -X POST "https://eu.api.foura.ai/api/browser/" \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/products", "proxy": "A1B2C3", "timeout_ms": 20000}'
ID 대신 주소를 전송하면 400 Invalid proxy format(으)로 반환됩니다. 전체 패턴은 Reuse a Proxy Across Requests에서 확인할 수 있습니다.
Next Steps
- Choosing the Right Endpoint: browser 및 single 사용 시점
- Monitor Competitor Prices: 전체 가격 추적 튜토리얼
- Protected sites: 보호된 사이트 처리 방법
- Reuse a Proxy Across Requests: 전체 워크플로에서 단일 출구 고정