Python crawler obtains images and downloads and saves them locally-Python Tutorial-php.cn

Home

Backend Development

Python Tutorial

Python crawler obtains images and downloads and saves them locally

不言

Jun 02, 2018 pm 02:50 PM

python download picture

This article mainly introduces the Python crawler to obtain images and download and save them locally. It has certain reference value. Now I share it with you. Friends in need can refer to it.

1. Grab Take the picture of omelette online.

2. The code is as follows:

import urllib.request
import os
#to open the url
def url_open(url):
 req=urllib.request.Request(url)
 req.add_header(&#39;User-Agent&#39;,&#39;Mozilla/5.0 (Windows NT 6.3; WOW64; rv:51.0) Gecko/20100101 Firefox/51.0&#39;)
 response=urllib.request.urlopen(url)
 html=response.read()
 return html
#to get the num of page like 1,2,3,4...
def get_page(url):
 html=url_open(url).decode(&#39;utf-8&#39;)
 a=html.find(&#39;current-comment-page&#39;)+23 #add the 23 offset th arrive at the [2356]
 b=html.find(&#39;]&#39;,a)
 #print(html[a:b])
 return html[a:b]
#find the url of imgs and return the url of arr
def find_imgs(url):
 html=url_open(url).decode(&#39;utf-8&#39;)
 img_addrs=[]
 a=html.find(&#39;img src=&#39;)
 while a!=-1:
  b=html.find(&#39;.jpg&#39;,a,a+255) # if false : return -1
  if b!=-1:
   img_addrs.append(&#39;http:&#39;+html[a+9:b+4])
  else:
   b=a+9
  a=html.find(&#39;img src=&#39;,b)
 #print(img_addrs)  
 return img_addrs
  #print(&#39;http:&#39;+each)
  
#save the imgs 
def save_imgs(folder,img_addrs):
 for each in img_addrs:
  filename=each.split(&#39;/&#39;)[-1] #get the last member of arr,that is the name
  with open(filename,&#39;wb&#39;) as f:
   img = url_open(each)
   f.write(img)
 
def download_mm(folder=&#39;mm&#39;,pages=10):
 os.mkdir(folder)
 os.chdir(folder)
 url=&#39;http://jandan.net/ooxx/&#39;
 page_num=int(get_page(url))
 
 for i in range(pages):
  page_num -= i
  page_url = url + &#39;page-&#39; + str(page_num) + &#39;#comments&#39;
  img_addrs=find_imgs(page_url)
  save_imgs(folder,img_addrs)
  
if __name__ == &#39;__main__&#39;:
 download_mm()

Copy after login

Related recommendations:

How to use Python crawler to get those valuable blog posts

Python crawler to get the website of American TV series

The above is the detailed content of Python crawler obtains images and downloads and saves them locally. For more information, please follow other related articles on the PHP Chinese website!

Statement of this Website

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Hot AI Tools

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress images for free

Clothoff.io

AI clothes remover

AI Hentai Generator

Generate AI Hentai for free.

Hot Article

R.E.P.O. Energy Crystals Explained and What They Do (Yellow Crystal)

2 weeks ago By 尊渡假赌尊渡假赌尊渡假赌

Repo: How To Revive Teammates

4 weeks ago By 尊渡假赌尊渡假赌尊渡假赌

Hello Kitty Island Adventure: How To Get Giant Seeds

4 weeks ago By 尊渡假赌尊渡假赌尊渡假赌

How Long Does It Take To Beat Split Fiction?

3 weeks ago By DDD

R.E.P.O. Save File Location: Where Is It & How to Protect It?

3 weeks ago By DDD

Hot Tools

Notepad++7.3.1

Easy-to-use and free code editor

SublimeText3 Chinese version

Chinese version, very easy to use

Zend Studio 13.0.1

Powerful PHP integrated development environment

Dreamweaver CS6

Visual web development tools

SublimeText3 Mac version

God-level code editing software (SublimeText3)

Hot Topics

Where is the login entrance for gmail email?

7322

Java Tutorial

1625

CakePHP Tutorial

1350

Laravel Tutorial

1262

PHP Tutorial

1209

Related knowledge

How to efficiently integrate Node.js or Python services under LAMP architecture? Apr 01, 2025 pm 02:48 PM

Many website developers face the problem of integrating Node.js or Python services under the LAMP architecture: the existing LAMP (Linux Apache MySQL PHP) architecture website needs...

How to solve the permissions problem encountered when viewing Python version in Linux terminal? Apr 01, 2025 pm 05:09 PM

Solution to permission issues when viewing Python version in Linux terminal When you try to view Python version in Linux terminal, enter python...

What is the reason why pipeline persistent storage files cannot be written when using Scapy crawler? Apr 01, 2025 pm 04:03 PM

When using Scapy crawler, the reason why pipeline persistent storage files cannot be written? Discussion When learning to use Scapy crawler for data crawler, you often encounter a...

Python hourglass graph drawing: How to avoid variable undefined errors? Apr 01, 2025 pm 06:27 PM

Getting started with Python: Hourglass Graphic Drawing and Input Verification This article will solve the variable definition problem encountered by a Python novice in the hourglass Graphic Drawing Program. Code...

What is the reason why the Python process pool handles concurrent TCP requests and causes the client to get stuck? Apr 01, 2025 pm 04:09 PM

Python process pool handles concurrent TCP requests that cause client to get stuck. When using Python for network programming, it is crucial to efficiently handle concurrent TCP requests. ...

How to view the original functions encapsulated internally by Python functools.partial object? Apr 01, 2025 pm 04:15 PM

Deeply explore the viewing method of Python functools.partial object in functools.partial using Python...

Python Cross-platform Desktop Application Development: Which GUI Library is the best for you? Apr 01, 2025 pm 05:24 PM

Choice of Python Cross-platform desktop application development library Many Python developers want to develop desktop applications that can run on both Windows and Linux systems...

Do Google and AWS provide public PyPI image sources? Apr 01, 2025 pm 05:15 PM

Many developers rely on PyPI (PythonPackageIndex)...

See all articles