Affichage des articles dont le libellé est Web. Afficher tous les articles
Affichage des articles dont le libellé est Web. Afficher tous les articles

jeudi 24 mai 2018

Advanced scraping

I wrote web scraping scripts lately for a client to download invoices. I really like web scraping because it’s an automation that allows to save a lot of time and it often demands to reverse engineer a site. Just like a real hacker, ma! Here are some techniques I use.

Disclaimer: using a robot to scrape a website can be prohibited by its owner, so check that you stay in the terms of usage conditions and please have a fair behaviour.

Tooling

Mostly, I use Python 3, Requests and BeautifulSoup. Everything can be installed with Pip.

Session

Send all your requests through a Session. It has several advantages:

  • it handles session cookies for you,
  • it keeps the TCP sockets open,
  • it allows you to set default header for all the requests you’ll do. For instance, to setup the user agent header in a session.

Like this:

# Wow, the UA is quite old!
session = requests.session()
session.headers.update(
    {
        "User-Agent": (
            "Mozilla/5.0 (X11; Linux x86_64; rv:58.0)"
            "Gecko/20100101 Firefox/58.0"
        )
    }
)

Authentication

When you want to gather data from the Internet, web sites will often require you need to sign in. Mostly, you’ll have to send your credentials within your session in a HTTP POST request:

session.post(your_url, data=form_data)

CSRF protection

A good practice when handling forms in a website is to associate them with a random token. Hence you have to load a the form from the website before submitting it. That’s basicly CSRF protection, to avoid a malicious site to make use of an existing session cookie to discretly perform post requests.

It’s a general good practice, though it’s questionnable on authentication form!

By the way, when authenticating, you’ll often need to load the sign-in page, gather all the input fields then update them with username and password. Something like this:

soup = BeautifulSoup(self._session.get(self._url).text, "html5lib")
login_data = {}
for input in soup.find_all("input"):
    if input["type"] in {"hidden", "text", "password"}:
        login_data[input["name"]] = input["value"]

login_data["credential"] = my_username
login_data["password"] = my_password

SAML

Sometimes, the authentication process is complicated. For a website, I needed to pass through a SAML authentication protocol.

SAML authentication consists in several exchanges between client and server to generate the authentication token. The token has to be extracted from the redirection history.

To do so, requests lib allows you to browse the redirection history from a response and update the session cookies.

In requests, access the response history with:

response.history  # get the response chain has a list

You can then get the cookies from one of the responses as a dict:

saml_response_cookies = requests.utils.dict_from_cookiejar(
    response.history[1].cookies
)

Finally, update your session cookie jar with these cookies:

session.cookies.update(saml_cookies)

The information is not in the source code

Classic web scraping works well when all the content of the page is sent in the html on page load. What if it is a Single Page App on which all the content is loaded dynamicly in JavaScript?

Look for XHR (use of AJAX) in the "Network" tab in the browser’s dev tools. You can then replay these XHR directly with requests and parse the response.

Most of the time, the response will be a JSON document. It’s good news, as parsing JSON content is far easier than parsing HTML. Just use json module from Python stdlib.

Sometimes the requests are built by the JavaScript code. I once scraped a website that converted session information in base64 to build a request QueryString in Javascript. By chance, these functions where readable, so I reimplemented them in my Python code.

I don’t see the whole document!

Sometimes, beautifulsoup silently fails to parse all your HTML document. It’s really annoying because you often spend a ridiculous amount of time to figure out that the error comes from beautiful soup parsing.

Try to vary parsers with beautifulsoup. Using html5lib gave me the best results. I think you can add this module directly with you dependencies and use it by default:

soup = BeautifulSoup(content, "html5lib")

When everything fails

Sometimes though, reverse engineering how a website works is too time consuming (JavaScript is minified, lot of state information is kept on client side,…). I then switch to headless browser automation. I do so with Chrome with headless option and Selenium.

This is a heavier approach as you need to script all the actions which would be performed by a real user: loading a page, waiting for all the content to be loaded, fill in form fields, click on buttons.

Usually, I fire up Chrome without the headless option during the script writing.I switch back to headless when everyting works.

samedi 7 avril 2018

Download files with Python, Selenium and Chrome headless

I have recurring tasks these days that consist to automatically download files on the Internet. Usually, requests, beautifulsoup and some tricks do the job effectively. Sometimes though, I have to play it hard and ask Selenium and Chromium* headless to do the heavy lifting. Alas, asking Chromium to automatically download files is not clear.

I found the solution in Chrome tracker, and you read it bellow written in Python:

from selenium import webdriver

# Let's create some option to make Chromium go headless
options = webdriver.ChromeOptions()
options.add_argument('headless')
options.add_argument('disable-gpu')

# Launch the browser 
browser = webdriver.Chrome(chrome_options=options)
download_dir = tempfile.TemporaryDirectory().name
os.mkdir(download_dir)

# Send a command to tell chrome to download files in download_dir without
# asking.
browser.command_executor._commands["send_command"] = (
    "POST",
    '/session/$sessionId/chromium/send_command'
)
params = {
    'cmd': 'Page.setDownloadBehavior',
    'params': {
        'behavior': 'allow',
        'downloadPath': download_dir
    }
}
browser.execute("send_command", params)

There you go, happy scraping!

* Of course, it works with regular Chrome too!

mercredi 21 octobre 2015

Back to Standards

I’m not so much a front-end guy. I mean, I do love hacking CSS and JS for the browser, but I don’t do it that much as a pro to be aware of all the best practices in this domain. And JS being the top class citizen in the Web this day, this field earned a huge complexity within a blink and I feel a bit overwhelmed. I’m still nicely editing my JS and CSS, refreshing the browser to see the results at the time of preprocessing, transpiling and minifying megabytes of code.

I often wonder if it doesn’t go too far sometimes, to the point that programmers need to reimplement a behavior brought for free by the browser. The must funny example is the use of the prev and next buttons in one page app. But you can also reimplement every form input elements.

That’s what I did, in a project I decided to replace input fields and buttons by plain divs. Because I found that contenteditable attribute was cool. I believed it’ll be easier to style everything.
I realized my mistake lately as getting all in place was a PITA. Moreover, now that I use vimperator, I found out that it would be awful to use the app without mouse if I did not recode the whole shit by myself. So I brought back my text areas and buttons tags. I just “unstyled” them and rework my JS a little bit. I was done in less than a hour.

So here’s the lesson. If you are not a professional front end developer and you need to get things done quickly, consider these points:

  • Keep it stupid simple. Really. Rely on what is offered by the browser. Do not try to be smart because this front end thing is waaaay too complicated for you.
  • It’s hard to admit this because I hate it… But use JQuery as it will eliminate a lot of burden. Coding in vanilla js when you have DOM manipulation is fun until you realize you spend hours to fix a simple UI behavior. And you’ll need few days more to make your fix cross browsers.
  • Really consider to use UI boilerplates like Foundation or Bootstrap. They’re blazingly simple to put in place, provide an elegant look and feel and make your site responsive for free. I know you see the same UI everywhere on the Internet. That’s because the other developers and startupers are like you: they lack skills and time and yet, need to get things done quickly.

This post raises the question of the choice between learning things deeply and get side projects done and be proud of them (where proud can mean getting money). It’ll be a great topic for another post.

lundi 25 février 2013

Frontend

Le monde m'en est témoin (parce que ça le concerne) : je fais actuellement du développement frontend (pendant mon temps libre, bien sûr). En HTML5 et JavaScript, s'il vous plait. Je trouve ça génial. C'est un peu un retour à mes premières amours car j'ai réellement appris le développement utile au travers du web avec PHP3. JavaScript avait alors mauvaise presse (nous étions en 2004 / 2005). Pour ma part, je trouvais ça puissant mais compliqué et peu élégant. Puissant car ça m'a permis de faire des trucs super cool à une époque encore antérieure quand j'avais utilisé Dreamweaver, c'est à dire sans écrire une seule ligne de code. Compliqué parce que quand j'essayais de coder, je ne comprenais pas la logique : ça ne marchais pas, le browser me crachait plein d'erreurs. Inélégant car l'usage de l'époque était de l'intégrer dans le HTML avec des attributs du genre onclick="code js dégueu". Cette façon de faire faisait un mélange des genres assez détestable !

Et puis j'ai découvert le JavaScript non intrusif et par ce biais qu'il s'agissait d'un vrai langage de haut niveau. Plus j'avance et plus je trouve ces particularités super sympathiques. Par exemple, l'aspect asynchrone : on ne peux pas faire de boucle d'attente en JS, il fait les appels et continue son exécution. Ça ouvre vers d'autres styles de programmation.

Je n'utilise pas de bibliothèque type Jquery (même si j'ai déjà testé, je vous rassure). Pour l'instant, je n'ai pas de problème avec le pur JS et j'aimerais continuer dans cette voie pour acquérir une meilleure maîtrise des idiomes du langage. Toutefois, je sais qu'en cas de besoin, je peux profiter d'un riche écosystème open source (ce qui permet également d'apprendre !)

JE joue également avec SVG. Ce qui est bien avec SVG c'est que comme c'est du XML, cela nous donne des images dont la définition est intégrée au DOM, et donc exploitable avec du JavaScript !

Pour conclure, je suis sur un projet assez sérieux avec tout ça. J'espère pouvoir le mener à terme. J'espère en tout cas qu'il saura trouver une utilité.

lundi 30 août 2010

Javascript

Il y a une époque, tout le monde considérait que le JavaScript, c'était sale et globalement pourri. Maintenant que Google est passé par là et que la navigation sur Internet nécessite un CPU 4 cœurs pour un minimum de confort, le paradigme a un peu changé.

Alors pour faire un site pour madame, je m'y mets.

Au départ à l'ancienne, avec les évènements imbriqués dans le HTML (pas taper, merci).

Puis j'ai découvert le concept de JavaScript non-intrusif.

Et là, ça commence à être classe !