Reason for choosing curl
Regarding curl and file_get_contents, here is an easy-to-understand comparison:
file_get_contents is actually a merged version of a bunch of built-in file operation functions, such as file_exists, fopen, fread, fclose, specially provided for lazy people. And it is mainly used to deal with local files, but because of lazy people, it also adds support for network files;
curl is a library specially used for network interaction, providing a bunch of custom options , used to deal with different environments, and its stability is naturally greater than file_get_contents.
How to use
1. Enable curl support
Since the curl support is not turned on by default after the PHP environment is installed, you need to modify the php.ini file, find; extension=php_curl.dll, remove the colon in front, and restart the service;
2. Use curl to capture data
Copy code The code is as follows:
//Initialize a cURL object
$curl = curl_init();
//Set the URL you need to crawl
curl_setopt($curl, CURLOPT_URL, 'http://www.cmx8.cn');
// Set header
curl_setopt($curl, CURLOPT_HEADER, 1);
//Set cURL parameters and require the results to be saved in a string or output to the screen.
curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1);
// Run cURL and request the web page
$data = curl_exec($curl);
// Close URL request
curl_close($curl );
3. Find key data through regular matching
Copy code The code is as follows:
//$data is the value returned by curl_exec, which is the target content collected
preg_match_all("/
(.*?)/",$data, $out, PREG_SET_ORDER);
foreach($out as $key => $value){
//Here $value is an array, and records the entire sentence with matching characters and the individually matched characters
echo 'The entire sentence matched: '.$value[0].'
';
echo 'Single match: '.$value[1].'
';
}
Tips
1. Timeout related settings
You can set some timeout settings through curl_setopt($ch, opt), mainly including:
CURLOPT_TIMEOUT sets the maximum number of seconds cURL is allowed to execute.
CURLOPT_TIMEOUT_MS sets the maximum number of milliseconds cURL is allowed to execute. (Added in cURL 7.16.2. Available from PHP 5.2.3.)
CURLOPT_CONNECTTIMEOUT The time to wait before initiating a connection. If set to 0, it will wait indefinitely.
CURLOPT_CONNECTTIMEOUT_MS The time to wait for a connection attempt, in milliseconds. If set to 0, wait infinitely. Added in cURL 7.16.2. Available starting with PHP 5.2.3.
CURLOPT_DNS_CACHE_TIMEOUT sets the time to save DNS information in memory, the default is 120 seconds.
Copy code The code is as follows:
curl_setopt($ch, CURLOPT_TIMEOUT, 60); //Only need to set one second The number can be
curl_setopt($ch, CURLOPT_NOSIGNAL, 1); //Note that the millisecond timeout must be set
curl_setopt($ch, CURLOPT_TIMEOUT_MS, 200); //The timeout in milliseconds is changed in cURL 7.16.2 join in. Available from PHP 5.2.3
2. Submit data through post and retain cookies
Copy code The code is as follows:
//The following is an example for learning and reference:
//Curl simulates login discuz program, suitable for DZ7.0
!extension_loaded('curl') && die( 'The curl extension is not loaded.');
$discuz_url = 'http://www.lxvoip.com';//Forum address
$login_url = $discuz_url .'/logging.php ?action=login'; //Login page address
$get_url = $discuz_url .'/my.php?item=threads'; //My post
$post_fields = array();
//The following two items do not need to be modified
$post_fields['loginfield'] = 'username';
$post_fields['loginsubmit'] = 'true';
//Username and password are required Fill in
$post_fields['username'] = 'lxvoip';
$post_fields['password'] = '88888888';
//Security question
$post_fields['questionid'] = 0 ;
$post_fields['answer'] = '';
//@todo verification code
$post_fields['seccoverify'] = ''; 🎜>$ch = curl_init($login_url);
curl_setopt($ch, CURLOPT_HEADER, 0);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1); 🎜>curl_close($ch);
preg_match('//i' , $contents, $matches);
if(!empty($matches)) {
$formhash = $matches[1];
} else {
die('Not found the forumhash. ');
}
//POST data, get COOKIE
$cookie_file = dirname(__FILE__) . '/cookie.txt';
//$cookie_file = tempnam('/ tmp');
$ch = curl_init($login_url);
curl_setopt($ch, CURLOPT_HEADER, 0);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1); CURLOPT_POST, 1);
curl_setopt($ch, CURLOPT_POSTFIELDS, $post_fields);
curl_setopt($ch, CURLOPT_COOKIEJAR, $cookie_file);
curl_exec($ch);
curl_close($ch) ;
//Use the COOKIE obtained above to obtain the content of the page that needs to be logged in to view.
$ch = curl_init($get_url);
curl_setopt($ch, CURLOPT_HEADER, 0);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, 0);
curl_setopt($ch, CURLOPT_COOKIEFILE, $cookie_file);
$contents = curl_exec($ch);
var_dump($contents);
http://www.bkjia.com/PHPjc/728088.html
www.bkjia.com
true
http: //www.bkjia.com/PHPjc/728088.html
TechArticleReasons for choosing curl Regarding curl and file_get_contents, here is an easy-to-understand comparison: file_get_contents is actually a bunch of built-in Merged versions of file operation functions, such as file_ex...