猫眼电影票房爬取到MySQL中_爬取猫眼电影top100，request、beautifulsoup运用

article/2025/11/10 5:15:03

这是第三篇爬虫实战，运用request请求，beautifulsoup解析，mysql储存。

如果你正在学习爬虫，本文是比较好的选择，建议在学习的时候打开猫眼电影top100进行标签的选择，具体分析步骤就省略啦，具体的方法可看我以前的总结。

练习猫眼top100是为了复习request和beautifulsoup以及sql中表的创建、字段类型的设置。好了附上源代码

import requests

from bs4 import BeautifulSoup

from requests.exceptions import RequestException

headers = {

'user-agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.102 Safari/537.36"}

def get_page_index(url):

try:

respones = requests.get(url, headers=headers) # get请求

if respones.status_code == 200:

return respones.text

return None

except RequestException:

print("请求有错误")

return None

def parse_one_page(html):

soup = BeautifulSoup(html, 'lxml')

board_wrapper = soup.find_all('dl', class_='board-wrapper')[0]

board_index = board_wrapper.find_all('dd')

for link in board_index:

title = link.find_all('p', class_='name')[0].text

board_index = link.find_all('i')[0].text

star = link.find_all('p', class_='star')[0].text.strip()[3:]

releasetime = link.find_all('p', class_='releasetime')[0].text.strip()[5:]

score = link.find_all('p', class_='score')[0].text

print(board_index, title, star, releasetime, score)

def main():

t = [0, 10, 20, 30, 40, 50, 60, 70, 80, 90]

urls = ['https://maoyan.com/board/4?offset={}'.format(i) for i in t]

for url in urls:

html = get_page_index(url)

parse_one_page(html)

if __name__ == "__main__":

main()

爬取内容如下:

存入数据库

存入数据库01：新建表

ALTER TABLE maoyantop100 CONVERT TO CHARACTER SET utf8 COLLATE utf8_unicode_ci;

create table maoyantop100

(board_index int not null,

title varchar(20) not null,

star varchar(50) not null,

releasetime varchar(50) not null,

score float not null);

存入数据库02

import pymysql

import requests

import json

from bs4 import BeautifulSoup

from requests.exceptions import RequestException

headers = {

'user-agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.102 Safari/537.36"}

def get_page_index(url):

try:

respones = requests.get(url, headers=headers) # get请求

if respones.status_code == 200:

return respones.text

return None

except RequestException:

print("请求有错误")

return None

def parse_one_page(html):

soup = BeautifulSoup(html, 'lxml')

board_wrapper = soup.find_all('dl', class_='board-wrapper')[0]

board_index = board_wrapper.find_all('dd')

for link in board_index:

title = link.find_all('p', class_='name')[0].text

board_index = link.find_all('i')[0].text

star = link.find_all('p', class_='star')[0].text.strip()[3:]

releasetime = link.find_all('p', class_='releasetime')[0].text.strip()[5:]

score = link.find_all('p', class_='score')[0].text

return board_index, title, star, releasetime, score

def save_to_mysql(board_index,title,star, releasetime,score):

cur = conn.cursor() #用来获得python执行Mysql命令的方法，操作游标。cur = conn.sursor才是真的执行

insert_data = "INSERT INTO maoyantop100(board_index,title,star, releasetime,score)" "VALUES(%s,%s,%s,%s,%s)"

val = (board_index,title,star, releasetime,score)

cur.execute(insert_data, val)

conn.commit()

conn=pymysql.connect(host="localhost",user="root", password='180****3940',db='58yuesao',port=3306, charset="utf8")

def main():

t = [0, 10, 20, 30, 40, 50, 60, 70, 80, 90]

urls = ['https://maoyan.com/board/4?offset={}'.format(i) for i in t]

for url in urls:

html = get_page_index(url)

board_index, title, star, releasetime, score = parse_one_page(html)

save_to_mysql(board_index, title, star, releasetime, score)

if __name__ == "__main__":

main()

存入数据库

本文存在一个小bug，在存入数据库时很崩溃，只爬取了每页的第一条信息。

我也不知道为什么，没存入数据库时爬取内容是完整的，存入之后就…………解决后会更新到本文。

作者：夜希辰

猫眼电影票房爬取到MySQL中_爬取猫眼电影top100，request、beautifulsoup运用

相关文章

python 抢票代码猫眼演出_Python爬虫-猫眼电影排行

flask+猫眼电影票房预测和电影推荐

猫眼电影产品分析

超过53亿！《长津湖》为什么这么火爆？我用 Python 来分析猫眼影评

python使用pyecharts对猫眼电影票房精美可视化分析简单仪表盘？？（五个图好多个组件！！）

python爬猫眼电影影评,Python系列爬虫之爬取并简单分析猫眼电影影评

爬取猫眼电影，进行分析

Python—猫眼电影票房爬虫实战轻松弄懂字体反爬！

基于Python猫眼票房TOP100电影数据抓取

基于猫眼票房数据的可视化分析

ardruino控制继电器_Arduino 各种模块篇-继电器

【继电器模块教程基于Arduino】

8路USB继电器模块 windows Linux使用

ardruino控制继电器_arduino控制继电器

Arduino笔记实验(初级阶段)—继电器模块

继电器模块典型电路图

固态继电器和电磁继电器模块

【继电器模块的电路设计和分析】

常用继电器模块的PCB设计与实物分享

4入4出Modbus RTU继电器模块说明书