The file_get_contents function is rarely used in batch data collection in PHP, but if it is a small amount, we can use the file_get_contents function, because it is not only easy to use but also easy to learn. Let me introduce the usage and use of file_get_contents Problem solving in the process.
Let’s look at the problem first
file_get_contents cannot obtain the URL with port
For example:
代码如下 | 复制代码 |
file_get_contents('http://localhost:12345'); |
Nothing gained.
The solution is: turn off selinux
1 Permanent method – need to restart the server
Modify the /etc/selinux/config file to set SELINUX=disabled, and then restart the server.
2 Temporary method – Set system parameters
Use command setenforce 0
Attachment:
setenforce 1 Set SELinux to enforcing mode
setenforce 0 Set SELinux to permissive mode
file_get_contents timeout
The code is as follows | Copy code | ||||
|
After the above problems are solved, we can start collecting.
代码如下 | 复制代码 |
//全国,判断条件是$REQUEST_URI是否含有html if (!strpos($_SERVER["REQUEST_URI"],".html")) { $page="http://qq.ip138.com/weather/"; $html = file_get_contents($page,'r'); $pattern="/全国主要城市、县当天和未来五天天气趋势预报在线查询(.*?) //正则匹配之间的html preg_match($pattern,$html,$pg); echo ""; //正则替换远程地址为本地地址 $p=preg_replace('//weather/(w+)/index.htm/', 'tq.php/.html', $pg[1]); echo $p; } //省,判断条件是$REQUEST_URI是否含有? else if(!strpos($_SERVER["REQUEST_URI"],"?")){ //yoyo推荐的使用分割获得数据,这里是获得省份名称 $province=explode("/",$_SERVER["REQUEST_URI"]); $province=explode(".",$province[count($province)-1]); $province=$province[0]; //被注释掉的是我自己写出来的正则,感觉写的不好,但效果等同上面 //preg_match('/[^/]+[.(html)]$/',$_SERVER["REQUEST_URI"],$pro); //$province=preg_replace('/.html/','',$pro[0]); $page="http://qq.ip138.com/weather/".$province."/index.htm"; //获取html数据之前先尝试打开页面,防止恶意输入地址导致出错 if (!@fopen($page, "r")) { die("对不起,该地址不存在!点击这里返回"); exit(0); } $html = file_get_contents($page,'r'); $pattern="/五天天气趋势预报(.*?)请输入输入市/si"; preg_match($pattern,$html,$pg); echo ""; //正则替换,获取省份,城市 $p=preg_replace('//weather/(w+)/(w+).htm/', '.html?pro=', $pg[1]); echo $p; } else { //市,通过get传递省份 $pro=$_REQUEST['pro']; $city=explode("/",$_SERVER["REQUEST_URI"]); $city=explode(".",$city[count($city)-1]); $city=$city[0]; //preg_match('/[^/]+[.(html)]+[?]/',$_SERVER["REQUEST_URI"],$cit); //$city=preg_replace('/.html?/','',$cit[0]); $page="http://qq.ip138.com/weather/".$pro."/".$city.".htm"; if (!@fopen($page, "r")) { die("对不起,该地址不存在!点击这里返回"); exit(0); } $html = file_get_contents($page,'r'); $pattern="/五天天气趋势预报(.*?)请输入输入市/si"; preg_match($pattern,$html,$pg); echo ""; //获取真实的图片地址 $p=preg_replace('//image//', 'http://qq.ip138.com/image/', $pg[1]); echo $p; } ?> |
If the above method cannot collect the data, we can use it to process it
代码如下 | 复制代码 |
$url = "http://www.bKjia.c0m "; $ch = curl_init(); $timeout = 5; curl_setopt($ch, CURLOPT_URL, $url); curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1); curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, $timeout); //在需要用户检测的网页里需要增加下面两行 //curl_setopt($ch, CURLOPT_HTTPAUTH, CURLAUTH_ANY); //curl_setopt($ch, CURLOPT_USERPWD, US_NAME.":".US_PWD); $contents = curl_exec($ch); curl_close($ch); echo $contents; ?> |