web scraping - Python BeautifulSoup - Scrape Web Content Inside Iframes

Question

Welcome To Ask or Share your Answers For Others

web scraping - Python BeautifulSoup - Scrape Web Content Inside Iframes

asked Oct 24, 2021 in Technique[技术] by 深蓝 (71.8m points)

web scraping - Python BeautifulSoup - Scrape Web Content Inside Iframes

We have this URL: https://www.aliexpress.com/store/feedback-score/1665279.html

And the needed content is the "Feedback History" table, which is inside an iframe:

Feedback    1 Month 3 Months    6 Months
Positive (4-5 Stars)    154 562 1,550
Neutral (3 Stars)   8   19  65
Negative (1-2 Stars)    8   20  57
Positive feedback rate  95.1%   96.6%   96.5%

How do we extract it?

See Question&Answers more detail:os

与恶龙缠斗过久,自身亦成为恶龙；凝视深渊过久,深渊将回以凝视…

1 Answer

深蓝 · Answer 1 · 2021-10-23T19:18:22+0000

You just need to obtain the src attribute of the iframe, and then request and parse its content:

import requests
from bs4 import BeautifulSoup

s = requests.Session()
r = s.get("https://www.aliexpress.com/store/feedback-score/1665279.html")

soup = BeautifulSoup(r.content, "html.parser")
iframe_src = soup.select_one("#detail-displayer").attrs["src"]

r = s.get(f"https:{iframe_src}")

soup = BeautifulSoup(r.content, "html.parser")
for row in soup.select(".history-tb tr"):
    print("".join([e.text for e in row.select("th, td")]))

Result:

Feedback        1 Month         3 Months        6 Months
Positive (4-5 Stars)    154     562     1,550
Neutral (3 Stars)       8       19      65
Negative (1-2 Stars)    8       20      57
Positive feedback rate  95.1%   96.6%   96.5%

Categories

web scraping - Python BeautifulSoup - Scrape Web Content Inside Iframes

web scraping - Python BeautifulSoup - Scrape Web Content Inside Iframes

Please log in or register to add a comment.

Please log in or register to answer this question.

1 Answer

Please log in or register to add a comment.

Just Browsing Browsing

Most popular tags