推荐学习书目
› Learn Python the Hard Way
Python Sites
› PyPI - Python Package Index
› http://diveintopython.org/toc/index.html
› Pocoo
值得关注的项目
› PyPy
› Celery
› Jinja2
› Read the Docs
› gevent
› pyenv
› virtualenv
› Stackless Python
› Beautiful Soup
› 结巴中文分词
› Green Unicorn
› Sentry
› Shovel
› Pyflakes
› pytest
Python 编程
› pep8 Checker
Styles
› PEP 8
› Google Python Style Guide
› Code Style from The Hitchhiker's Guide
qazwsxkevin
V2EX  ›  Python

请教这种网页的部分内容, Python 如何爬? curl&wget 获取的静态 html 获取不到这部分的内容。。。

  •  
  •   qazwsxkevin · Jan 25, 2019 · 2608 views
    This topic created in 2806 days ago, the information mentioned may be changed or developed.
    我不 duqiu 的,这个页面就遇到以往学习中没遇到过的情况。。
    如:
    http://www.310win.com/analysis/1659945.htm ,昨晚中国对伊朗的情况:
    在
    “对赛往绩”
    “中国 近期战绩”
    “伊朗 近期战绩”

    网页上有其中三个这样的数据表格,关于表格的内容:

    1、使用 cur 或者 wget 去获取 http://www.310win.com/analysis/1659945.htm ,默认获取的网页明文,是没有这样设计的三个表格的内容数据的。。。
    2、同上,页面静态打开,“对赛往绩”是有“平均欧赔”,“竞赛让球”“竞赛胜平负”等等下拉菜单选项等等东西,默认获取的网页明文,连这些下拉菜单的菜单内容都没有。。。
    3、第一个问题:python 如何获取这些内容?
    4、第二个问题,如果不确定表格下拉菜单有多少个(也许有可能根据不同的页面,有不同数量的菜单选择),python 如何逐步穷尽选择下拉菜单每一个,获取到每一个菜单选项都出现的内容?
    3 replies  •  2019-01-26 11:14:51 +08:00
    cxtrinityy
        1
    cxtrinityy  
       Jan 25, 2019
    1、2、3 是一个问题,你 wget 是下载的 html 初始化页面,你想拉取的是 js 渲染完后的 html,所以你下载和你浏览器里看到的页面内容不同
    所以简单说,是你的思路方向错了,你去网上搜一下如何用 python 获取 js 渲染的网页内容就行,overflow 上应该有相关资料
    580a388da131
        2
    580a388da131  
       Jan 25, 2019
    Phantomjs 渲染一次
    Ghosin
        3
    Ghosin  
       Jan 26, 2019
    from selenium import webdriver
    About   ·   Help   ·   Advertise   ·   Blog   ·   API   ·   FAQ   ·   Privacy   ·   Solana   ·   1149 Online   Highest 6679   ·     Select Language
    创意工作者们的社区
    World is powered by solitude
    VERSION: 3.9.8.5 · 27ms · UTC 17:13 · PVG 01:13 · LAX 10:13 · JFK 13:13
    ♥ Do have faith in what you're doing.