《Python + Selenium 自动登录抓取入门》

《Python + Selenium 自动登录抓取入门》

1 简介

Selenium 通过 WebDriver 驱动真实浏览器完成页面自动化,常用于接口冒烟、UI 回归与需要登录态的抓取任务。Python 绑定 + ChromeDriver 是最常见的组合。

2 环境安装

2.1 安装浏览器与驱动

Debian 安装 Google Chrome:

wget -q -O - https://dl-ssl.google.com/linux/linux_signing_key.pub | apt-key add - echo "deb http://dl.google.com/linux/deb/ stable main" >> /etc/apt/sources.list apt-get update apt-get install -y google-chrome-stable

安装 chromedriver 与 selenium:

apt-get install -y chromedriver python-pip pip3 install selenium

chromedriver 版本需与 Chrome 主版本一致,可从 http://chromedriver.chromium.org/ 下载匹配版本并放到 /usr/local/bin。

2.2 初始化无头浏览器

from selenium import webdriver def init_driver(driver_path="/usr/local/bin/chromedriver"): option = webdriver.ChromeOptions() option.add_argument("headless") # 无头模式,服务器上无需桌面 option.add_argument("--no-sandbox") driver = webdriver.Chrome(executable_path=driver_path, options=option) driver.implicitly_wait(10) return driver

3 定位元素与登录操作

Selenium 支持多种锚定方式:id、name、class、XPath、CSS 等。用 XPath 定位输入框并输入账号密码:

from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def login(driver, url, account, password): driver.get(url) # 定位用户名输入框(以 GitHub 登录页为例) input_user = driver.find_element_by_xpath("//*[@id=\"login_field\"]") input_user.clear() input_user.send_keys(account) driver.find_element_by_xpath("//*[@id=\"password\"]").send_keys(password) # 点击登录按钮 driver.find_element_by_xpath("//*[@id=\"login\"]/form/div[3]/input[3]").click()

4 显式等待

网络慢时元素可能尚未渲染,用显式等待代替 sleep:

def click_profile_menu(driver): driver.find_element_by_xpath('//*[@id="user-links"]/li[3]/details/summary/img').click() wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located( (By.XPATH, '//*[@id="user-links"]/li[3]/details/details-menu/ul/li[3]/a') )).click()

5 下载页面资源

取元素的 src 属性并下载:

from urllib import request def download_image(driver, css_selector, save_as="avatar.jpg"): img = driver.find_element_by_css_selector(css_selector) img_src = img.get_attribute("src") request.urlretrieve(img_src, save_as)

6 完整脚本

#!/usr/bin/python3 # -*- coding: utf-8 -*- import sys from urllib import request from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait URL = "https://example.com/login" ACCOUNT = sys.argv[1] if len(sys.argv) > 1 else "" PASSWORD = sys.argv[2] if len(sys.argv) > 2 else "" def init_driver(path): option = webdriver.ChromeOptions() option.add_argument("headless") return webdriver.Chrome(executable_path=path, options=option) def do_job(driver): driver.get(URL) driver.find_element_by_xpath("//input[@id='login_field']").send_keys(ACCOUNT) driver.find_element_by_xpath("//input[@id='password']").send_keys(PASSWORD) driver.find_element_by_xpath("//input[@type='submit']").click() wait = WebDriverWait(driver, 10) img = wait.until(EC.presence_of_element_located( (By.CSS_SELECTOR, "img.avatar"))).get_attribute("src") request.urlretrieve(img, "avatar.jpg") if __name__ == "__main__": driver = init_driver("/usr/local/bin/chromedriver") try: do_job(driver) finally: driver.quit()

7 常见问题

7.1 元素可定位但点击无效

  • 确认页面有多个同名元素,改用更精确的 XPath / CSS 或 find_elements 取目标;
  • 元素被遮挡时先 driver.execute_script("arguments[0].click();", el)。

7.2 找不到 chromedriver

  • 将 chromedriver 放到 PATH 或显式传入 executable_path;
  • 确保其与本地 Chrome 主版本一致,否则启动即报 session not created。

8 学习资料

阅读 — · 全站 —
🎸 我的歌单 0 首