php网络爬虫技术

时间:2014-07-22 14:51来源: 作者: 点击: 次

分享到：

<无详细内容>

function get_urls($url){  

       $url_array=array();  

       $the_first_content=file_get_contents($url);  

       $the_second_content=file_get_contents($url);  

       $pattern1 = "/http:\/\/[a-zA-Z0-9\.\?\/\-\=\&\:\+\-\_\'\"]+/";  

       $pattern2="/http:\/\/[a-zA-Z0-9\.]+/";  

       preg_match_all($pattern2, $the_second_content, $matches2);  

       preg_match_all($pattern1, $the_first_content, $matches1);  

       $new_array1=array_unique($matches1[0]);  

       $new_array2=array_unique($matches2[0]);  

       $final_array=array_merge($new_array1,$new_array2);  

       $final_array=array_unique($final_array);  

       for($i=0;$i<count($final_array);$i++)  

       {  

          echo $final_array[$i]."<br/>";  

       }  

   }  

    get_urls("http://www.baidu.com");

分享到： QQ空间新浪微博人人网开心网更多

精彩图集

Sublime里直接

Laravel 4 初级

PHP仿博客园

基于php中使

将word转化为

精彩文章

热点文章

php网络爬虫技术

热门标签

赞助商链接