Home Backend Development PHP Tutorial Use Crawler component to analyze HTML in laravel

Use Crawler component to analyze HTML in laravel

Aug 07, 2017 pm 05:10 PM
html laravel

This article mainly introduces the use of Symfony's Crawler component to analyze HTML in laravel. Friends in need can refer to it

The full name of Crawler is DomCrawler, which is a component of the Symfony framework. What is outrageous is that DomCrawler does not have Chinese documentation, and Symfony has not translated this part, so development using DomCrawler can only be explored bit by bit. Now I will summarize the experience in the use process.

The first thing is to install


composer require symfony/dom-crawler
composer require symfony/css-selector
Copy after login

css-seelctor is a css selector. Some functions will be used when selecting nodes with css

The example used in the manual is


use Symfony\Component\DomCrawler\Crawler;
$html = <<<‘HTML‘
Hello World!
Hello Crawler!
HTML;
$crawler = new Crawler($html);
foreach ($crawler as $domElement)
{
var_dump($domElement->nodeName);
}
Copy after login

The printed result is


string ‘html‘ (length=4)
Copy after login

because of this paragraph The nodeName of the html code is html. My English is not good. When I started using it, I thought the program was wrong. . .

In the actual use process, if new Crawler ($html) will have garbled characters, it should be related to the page encoding, so you can use the following method, first initialize the crawler, and then add node


$crawler = new Crawler();
$crawler->addHtmlContent($html);
Copy after login

The second parameter of addHtmlContent is charset, and the default is utf-8.

For other examples, please refer to the official documentation, http://symfony.com/doc/current/components/dom_crawler.html

Record the usages that you have tried little by little at work

filterXPath(string $xpath) method, according to the manual, the parameter of this method is $xpath, and p, p and other blocks are often used.


echo $crawler->filterXPath(‘//body/p‘)->text();
echo $crawler->filterXPath(‘//body/p‘)->last()->text();
Copy after login

The output is the text of the first and next p tag block


var_dump($crawler->filterXPath(‘//body‘)->html());
Copy after login

Output the html within the body


foreach ($crawler->filterXPath(‘//body/p‘) as $i => $node) {
$c = new Crawler($node);
echo $c->filter(‘p‘)->text();
}
Copy after login

filterXPath obtains an array of DOMElement blocks. Each DOMElement block can use a new crawler object to continue parsing


$nodeValues =
$crawler->filterXPath(‘//body/p‘)->each(function (Crawler $node, $i) {
return $node->text();
});
Copy after login

crawler provides each loop and uses closure functions to simplify the code. However, please note that this way of writing $nodeValues ​​results in an array, which requires further processing.

Other usage


##

echo $crawler->filterXPath(‘//body/p‘)->attr(‘class‘);
Copy after login

You can get the value "message" of the class attribute corresponding to the first p tag ”


$crawler->filterXPath(‘//p[@class="样式"]‘)->filter(‘a‘)->attr(‘href‘);
$crawler->filterXPath(‘//p[@class="样式"]‘)->filter(‘a>img‘)->extract(array(‘alt‘, ‘href‘))
Copy after login
The above are some methods of obtaining tag attributes

Filter and filterXPath are different. The manual says css selector. I don’t quite understand. I It is understood that the elements contained in XPath nodes such as p need to be tried in actual development.

Generally speaking, I feel that DomCrawler is easier to use than simple html dom. Maybe it is because I use it more easily.

The above are just the basic functions of Crawler. For more usage, please refer to the functions in the Crawler part of the symfony manual

http://api.symfony.com/3.2/Symfony/Component/DomCrawler/Crawler .html

The main problem with Crawler is that there are too few examples. There are no usage examples in the function manual, so you can only explore it in actual use. . . .

symfony's documentation about DomCrawler, which contains a few examples

http://symfony.com/doc/current/components/dom_crawler.html

The above is the detailed content of Use Crawler component to analyze HTML in laravel. For more information, please follow other related articles on the PHP Chinese website!

Statement of this Website
The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Hot AI Tools

Undresser.AI Undress

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

AI Clothes Remover

Online AI tool for removing clothes from photos.

Undress AI Tool

Undress AI Tool

Undress images for free

Clothoff.io

Clothoff.io

AI clothes remover

AI Hentai Generator

AI Hentai Generator

Generate AI Hentai for free.

Hot Article

R.E.P.O. Energy Crystals Explained and What They Do (Yellow Crystal)
2 weeks ago By 尊渡假赌尊渡假赌尊渡假赌
Hello Kitty Island Adventure: How To Get Giant Seeds
1 months ago By 尊渡假赌尊渡假赌尊渡假赌
Two Point Museum: All Exhibits And Where To Find Them
1 months ago By 尊渡假赌尊渡假赌尊渡假赌

Hot Tools

Notepad++7.3.1

Notepad++7.3.1

Easy-to-use and free code editor

SublimeText3 Chinese version

SublimeText3 Chinese version

Chinese version, very easy to use

Zend Studio 13.0.1

Zend Studio 13.0.1

Powerful PHP integrated development environment

Dreamweaver CS6

Dreamweaver CS6

Visual web development tools

SublimeText3 Mac version

SublimeText3 Mac version

God-level code editing software (SublimeText3)

Nested Table in HTML Nested Table in HTML Sep 04, 2024 pm 04:49 PM

This is a guide to Nested Table in HTML. Here we discuss how to create a table within the table along with the respective examples.

Table Border in HTML Table Border in HTML Sep 04, 2024 pm 04:49 PM

Guide to Table Border in HTML. Here we discuss multiple ways for defining table-border with examples of the Table Border in HTML.

HTML margin-left HTML margin-left Sep 04, 2024 pm 04:48 PM

Guide to HTML margin-left. Here we discuss a brief overview on HTML margin-left and its Examples along with its Code Implementation.

HTML Table Layout HTML Table Layout Sep 04, 2024 pm 04:54 PM

Guide to HTML Table Layout. Here we discuss the Values of HTML Table Layout along with the examples and outputs n detail.

HTML Ordered List HTML Ordered List Sep 04, 2024 pm 04:43 PM

Guide to the HTML Ordered List. Here we also discuss introduction of HTML Ordered list and types along with their example respectively

How do you parse and process HTML/XML in PHP? How do you parse and process HTML/XML in PHP? Feb 07, 2025 am 11:57 AM

This tutorial demonstrates how to efficiently process XML documents using PHP. XML (eXtensible Markup Language) is a versatile text-based markup language designed for both human readability and machine parsing. It's commonly used for data storage an

Moving Text in HTML Moving Text in HTML Sep 04, 2024 pm 04:45 PM

Guide to Moving Text in HTML. Here we discuss an introduction, how marquee tag work with syntax and examples to implement.

HTML Input Placeholder HTML Input Placeholder Sep 04, 2024 pm 04:54 PM

Guide to HTML Input Placeholder. Here we discuss the Examples of HTML Input Placeholder along with the codes and outputs.

See all articles