使用requests库和urlretrieve下载pdf文件

moboyou 2025-06-13 07:56 78 浏览

一、代码如下：

import requests   #导入请求库
from urllib.request import urlretrieve     #从urllib.request导入下载函数urlretrieve
import re,time   #导入正则库和时间库
from lxml import etree   #从lxml导入etree类
def gethtml():   #定义函数gethtml用来下载pdf文件
    url="http://www.gov.cn/zhengce/pdfFile/downloadFile.htm"   #设置请求网址
    headers={
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/69.0.3497.100 Safari/537.36"
    }                 #设置请求头headers
    response=requests.get(url,headers=headers)    #通过headers伪装对网站url进行get请求，并将响应内容赋值给response变量
    response.encoding=response.apparent_encoding   #根据网页内容解析出网页的编码格式并赋值给响应的编码变量response.encoding
    html=response.text   #将网页的相应的文本内容赋值给html
    html=etree.HTML(html)   #对html构造了一个XPath解析对象并对自动修正并赋值给html
    result=html.xpath('//tbody/tr')   #使用xpath找到tr标签并赋值给result
    urllist=[]   #定义接收网址的空列表urllist
    for info in result:    #遍历result里的变量info
        try:   #尝试操作
            urllist.append("http://www.gov.cn"+info.xpath('./td[2]/a/@href')[-1])   #将解析到的td标签的href属性值的最后一个元素与"http://www.gov.cn"相加并添加到列表urllist中
        except:   #当接收到错误时，
            continue  #继续执行
    # print(urllist)
    for downurl in urllist:   #遍历urllist列表中的网址downurl
        urlretrieve(downurl,"E://IT/PYthon/PYTHON试验/gov/"+downurl.split("/")[-1])   #下载网址downurl，并保存到本机的E://IT/PYthon/PYTHON试验/gov/文件夹下面，文件名用下载网址的最后切割的名称
        print("E://IT/PYthon/PYTHON试验/gov/"+downurl.split("/")[-1]+"下载成功")   #打印下载成功
        time.sleep(0.1)   #每执行一次下载休眠0.1秒
gethtml()   #调用gethtml函数

二、代码运行结果如下：

E://IT/PYthon/PYTHON试验/gov/PDF_ALL.zip下载成功

E://IT/PYthon/PYTHON试验/gov/2020_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2019_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2018_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2017_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2016_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2015_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2014_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2013_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2012_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2011_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2010_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2009_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2008_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2007_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2006_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2005_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2004_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2003_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2002_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2001_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/2000_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1999_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1998_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1997_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1996_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1995_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1994muLu.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1994_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1993_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1992muLu.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1992_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1991_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1990_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1989_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1988_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1987_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1986_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1985_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1984_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1983_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1982_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1981_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1980_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1979_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1978_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1973_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1971_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1970_PDF.pdf下载成功

E://IT/PYthon/PYTHON试验/gov/1969_PDF.pdf下载成功

三、代码和代码运行结果如下图所示：

最终保存到本机的数据如下图所示：

css教程pdf下载

上一篇：Centos安装wkhtmltopdf，curl访问网页生成pdf文件
下一篇：网页转pdf，这个工具真好用

使用requests库和urlretrieve下载pdf文件

相关推荐

Excel技巧:SHEETSNA函数一键提取所有工作表名称批量生产目录

Excel HOUR函数:“小时”提取器_excel+hour函数提取器怎么用

关于Excel(WPS表格)中公式，可以从12个方面理解，学后无忧!

Excel(WPS表格)Tocol函数应用技巧案例解读，建议收藏备用!

Filter+Search信息管理不再难|多条件|模糊查找|Excel函数应用

FILTER函数介绍及经典用法12:FILTER+切片器的应用

WPS/Excel职场办公最常用的60个函数大全(含卡片)，效率翻倍!

收藏|查找神器Xlookup全集|一篇就够|Excel函数|图解教程

比Vlookup更好用的万金油查找技巧，用过的都说好，建议收藏备用

查找匹配，Vlookup函数公式，1分钟入门至精通!