Python爬虫入门案例6：scrapy的基本语法+使用scrapy进行网站数据爬取

news2026/2/11 19:00:11

几天前在本地终端使用pip下载scrapy遇到了很多麻烦，总是报错，花了很长时间都没有解决，最后发现pycharm里面自带终端！（狂喜），于是直接在pycharm终端里面写scrapy了

这样的好处就是每次不用切换路径了，pycharm会直接把路径定位到项目包的路径下，非常方便。

而且下载scrapy可以直接在一个文件里面写import scrapy，然后install scrapy包就可以了，很快就下完了。

这时候我们就可以直接进行scrapy程序的创建了。

基本语法：

（1）创建scrapy爬虫项目

scrapy startproject 项目名

（2）创建爬虫文件

scrapy genspider 爬虫文件名爬取的网页

（3）运行爬虫代码

scrapy crawl 爬虫的名字

这里的爬虫主代码，需要在spiders文件中写

下面举个例子，使用scrapy来爬取汽车之家的汽车型号，与其对应的价格

import scrapy


class CarsSpider(scrapy.Spider):
    name = "cars"
    allowed_domains = ["https://car.autohome.com.cn/price/brand-15.html"]
    start_urls = ["https://car.autohome.com.cn/price/brand-15.html"]

    def parse(self, response):
        print("-------------")
        name_list=response.xpath("//div[@class='main-title']/a/text()")
        price_list=response.xpath("//div[@class='main-lever']//span/span/text()")
        for i in range(len(name_list)):
            name=name_list[i].extract()
            price=price_list[i].extract()
            print("-------------")
            print(name,price)

爬取结果：