Home > Backend Development > Python Tutorial > Python handles crawling Chinese encoding and judging encoding

Python handles crawling Chinese encoding and judging encoding

高洛峰
Release: 2016-10-19 11:45:20
Original
1262 people have browsed it

In the process of developing my own crawler, some web pages are utf-8, some are gb2312, and some are gbk. If not processed, the collected data will be garbled. The solution is to process the html into a unified utf-8 encoding

Version python2.7

#coding:utf-8
import chardet
#抓取网页html
line = "http://www.pythontab.com"
html_1 = urllib2.urlopen(line,timeout=120).read()
encoding_dict = chardet.detect(html_1)
print encoding
web_encoding = encoding_dict['encoding']
#处理,整个html就不会是乱码。
if web_encoding == 'utf-8' or web_encoding == 'UTF-8':
html = html_1
else :
html = html_1.decode('gbk','ignore').encode('utf-8')
Copy after login


Related labels:
source:php.cn
Statement of this Website
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn
Popular Tutorials
More>
Latest Downloads
More>
Web Effects
Website Source Code
Website Materials
Front End Template