PHP:preg_replace_callback匹配中文的问题
代码:
<code>$html = preg_replace_callback("/(?<chinese>[\x{4e00}-\x{9fa5}]+)/u",array("self","wyc_chinese"),$html); ... 省略 ... public function wyc_chinese($matches) { return $matches['chinese'].'(Chinese)'; } </chinese></code>
问题:
$html为要提取的网页数据
如果$html是utf8编码的,则以上代码能正常执行(即能正常提取中文),但如果是其他编码的,则没法正常执行(无法匹配到汉字)
使用iconv转换$html的编码格式,也无法正常提取中文。
回复内容:
代码:
<code>$html = preg_replace_callback("/(?<chinese>[\x{4e00}-\x{9fa5}]+)/u",array("self","wyc_chinese"),$html); ... 省略 ... public function wyc_chinese($matches) { return $matches['chinese'].'(Chinese)'; } </chinese></code>
问题:
$html为要提取的网页数据
如果$html是utf8编码的,则以上代码能正常执行(即能正常提取中文),但如果是其他编码的,则没法正常执行(无法匹配到汉字)
使用iconv转换$html的编码格式,也无法正常提取中文。
以<meta charset="utf-8">
来识别编码是错误的.有些网页没有写meta,对于现代浏览器也会正常显示的(IE6有问题,IE7,IE8没测~)
应该根据HTTP响应头Content-Type: text/html; charset=UTF-8
来判断.如果没有返回charset
,就根据内容来自行判断了..
为了方便,最好将html转换为UTF-8
来进行正则匹配.
<?php //编辑器的编码格式为UTF-8(无BOM) $remote_url = 'http://segmentfault.com/q/1010000000450422'; $context = stream_context_create([ 'http' => [ 'method' => 'GET', ], ]); $html = file_get_contents($remote_url, false, $context); $html_encoding = mb_detect_encoding($html, ['UTF-8', 'CP936', 'ASCII']); //转换为UTF-8 $target_encoding = 'UTF-8'; $html = $target_encoding === $html_encoding ? $html : mb_convert_encoding($html, $target_encoding, $html_encoding); //匹配 $count = preg_match_all('#[\x{4e00}-\x{9fa5}]+#u', $html, $matches); var_dump($matches);
你这问题的核心是网页编码转换成UTF-8
你说源编码是"根据meta标签的charset字段来判断的"
我也是这样子做的, 不过我成功.
你没给出详尽代码,我不知道是你的代码哪里出错了,还是纯粹是我的人品比你好.
<code>require_once(__DIR__.'/wp-config.php'); $resp = wp_remote_get('http://51nb.com/'); $html = $resp['body']; preg_match('@charset=([-a-z0-9_]+)@i',$html,$charset); $html = iconv(strtoupper($charset[1]), "UTF-8", $html); preg_match_all("@\p{Han}+@u",$html,$m); echo '<meta charset="UTF-8" />'; print_r($m); exit; </code>
使用以上代码的iconv
不使用以上代码的iconv

Hot AI Tools

Undresser.AI Undress
AI-powered app for creating realistic nude photos

AI Clothes Remover
Online AI tool for removing clothes from photos.

Undress AI Tool
Undress images for free

Clothoff.io
AI clothes remover

AI Hentai Generator
Generate AI Hentai for free.

Hot Article

Hot Tools

Notepad++7.3.1
Easy-to-use and free code editor

SublimeText3 Chinese version
Chinese version, very easy to use

Zend Studio 13.0.1
Powerful PHP integrated development environment

Dreamweaver CS6
Visual web development tools

SublimeText3 Mac version
God-level code editing software (SublimeText3)

Hot Topics



PHP 8.4 brings several new features, security improvements, and performance improvements with healthy amounts of feature deprecations and removals. This guide explains how to install PHP 8.4 or upgrade to PHP 8.4 on Ubuntu, Debian, or their derivati

To work with date and time in cakephp4, we are going to make use of the available FrozenTime class.

CakePHP is an open-source framework for PHP. It is intended to make developing, deploying and maintaining applications much easier. CakePHP is based on a MVC-like architecture that is both powerful and easy to grasp. Models, Views, and Controllers gu

To work on file upload we are going to use the form helper. Here, is an example for file upload.

Validator can be created by adding the following two lines in the controller.

Visual Studio Code, also known as VS Code, is a free source code editor — or integrated development environment (IDE) — available for all major operating systems. With a large collection of extensions for many programming languages, VS Code can be c

CakePHP is an open source MVC framework. It makes developing, deploying and maintaining applications much easier. CakePHP has a number of libraries to reduce the overload of most common tasks.

This tutorial demonstrates how to efficiently process XML documents using PHP. XML (eXtensible Markup Language) is a versatile text-based markup language designed for both human readability and machine parsing. It's commonly used for data storage an
