Lighthouse has a new layout. Prefer the old one? Return to the old layout, and switch back any time from the link at the top of each page.

Activity (RSS)

Friday, November 18 2011
36 skip_links_like example was updated
35 New to Ruby: Need basic Toturial on How to use anemone to crawl a site? was updated
  • Alex Johnson
    Alex Johnson commented at 2:54 AM

    learning new markups is harrowing

       rvm pkg install openssl
       rvm remove 1.9.2 
       rvm install 1.9.2 --with-openssl-dir=$HOME/.rvm/usr
    
  • Alex Johnson
    Alex Johnson commented at 2:52 AM

    correction -- @@@ rvm pkg install openssl

    rvm remove 1.9.2 
    rvm install 1.9.2 --with-openssl-dir=$HOME/.rvm/usr
    
    
    
  • Alex Johnson
    Alex Johnson commented at 2:50 AM

    The error is due to ruby being installed without openssl support. To resolve this problem, you will have to remove ruby and install it again. In case you are us...

23 Allow root domain redirects was updated
  • Alex Johnson
    • Milestone order changed from 0 to 0
    by Alex Johnson at 1:51 AM

    In case you or someone else is still facing this problem, the following solution can be applied.

    Anemone.crawl("http://heise.de") do |anemone|
    anemone.on_every_...

Wednesday, September 07 2011
36 skip_links_like example was created
  • Ben
    Ben created the ticket at 10:55 AM

    Can you provide a skip_links_like example.

    I am having problems understanding the syntax. Thanks.

Thursday, August 18 2011
35 New to Ruby: Need basic Toturial on How to use anemone to crawl a site? was created
  • Sidra
    Sidra created the ticket at 11:25 PM

    Hi,
    I am new to Ruby (learning Stage) and am using Anemone for crawling a site.
    1) I have use command " gem install anemone" to install anemone on mac and recei...

Tuesday, September 07 2010
15 Canonicalize URLs was updated
  • Lee Hambley
    • Milestone order changed from 0 to 0
    by Lee Hambley at 9:53 AM

    Agreed completely, if there's a canonical URI/L specified in the source, it should be used (or at least made available).

Wednesday, September 01 2010
34 Unique URL behaviour and 301 was updated
Friday, August 06 2010
34 Unique URL behaviour and 301 was created
  • mtkd
    mtkd created the ticket at 5:35 AM

    If page A 301 redirects to page B, Anemone will try to crawl page B again if link occurs later.

Friday, July 30 2010
29 Double encoding links was updated
  • chris (at chriskite)
    • Assigned user set to chris (at chriskite)
    • State changed from new to open
    • Milestone set to 0.4.1
    • Milestone order changed from 0 to 0
    by chris (at chriskite) at 7:36 PM
31 Unique URL behaviour if Tokyo Cabinet directory incorrect was updated
  • chris (at chriskite)
    • Milestone set to 0.4.1
    • State changed from new to open
    • Assigned user set to chris (at chriskite)
    • Milestone order changed from 185604 to 0
    by chris (at chriskite) at 7:34 PM
30 No response code or referer after running for hours was updated
  • chris (at chriskite)
    chris (at chriskite) commented at 7:34 PM

    Any idea what the memory usage was like on your system towards the end of the crawl? Which storage engine are you using, the default in-memory hash or TokyoCabi...

  • chris (at chriskite)
    • State changed from new to open
    by chris (at chriskite) at 7:33 PM
26 Illegal instruction was updated
  • chris (at chriskite)
    • Milestone set to 0.4.1
    • Milestone order changed from 185601 to 0
    by chris (at chriskite) at 7:32 PM

    Do either of you guys have an example of a URL that causes this? I can't reproduce it with the sites I usually crawl.

  • chris (at chriskite)
    • State changed from new to open
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 7:30 PM
33 Timeouts bubble up through rescue blocks and kill tentacles was updated
  • chris (at chriskite)
    chris (at chriskite) commented at 7:32 PM

    Great catch, I'll include a fix for this in 0.4.1

  • chris (at chriskite)
    • State changed from new to open
    • Milestone set to 0.4.1
    • Milestone order changed from 185603 to 0
    by chris (at chriskite) at 7:31 PM
  • michael.harrington
    michael.harrington created the ticket at 12:39 PM

    The thread body of the tentacle looks like this:

        def run
          loop do
            link, referer, depth = @link_queue.deq
    
            break if link == :END
    
         ...
0.4.1 was created by chris (at chriskite) at 7:31 PM
 
4 of 4 tickets remain in this milestone
Thursday, July 29 2010
26 Illegal instruction was updated
  • michael.harrington
    • Milestone order changed from 0 to 0
    by michael.harrington at 10:06 AM

    I'm having this same issue.

30 No response code or referer after running for hours was updated
  • michael.harrington
    michael.harrington commented at 10:05 AM

    I changed my crawl to use 2 threads and discard page bodies, which appears to avoid this issue.

32 nofollow behaviour was updated
  • mtkd
    mtkd commented at 8:16 AM

    example

  • mtkd
    mtkd created the ticket at 8:10 AM

    Anemone appears to respect nofollow links if it is before the href, but not if it appears later:

    Ignores link

    Still follows link

31 Unique URL behaviour if Tokyo Cabinet directory incorrect was created
  • mtkd
    mtkd created the ticket at 8:05 AM

    Minor issue, when using Tokyo Cabinet the URLs returned in a .on_every_page loop stop being unique quickly if the directory specified for the database file is i...

Monday, July 26 2010
30 No response code or referer after running for hours was created
  • michael.harrington
    michael.harrington created the ticket at 1:09 PM

    OS X Snow Leopard
    Ruby 1.8.7 (System Ruby)

    require 'rubygems'
    require 'bundler'
    Bundler.setup
    require 'anemone'
    
    files = {}
    
    Anemone.crawl 'http://local.acton....
Friday, June 25 2010
29 Double encoding links was created
  • Ben VandenBos
    Ben VandenBos created the ticket at 1:59 PM

    It seems that links with %20's get double encoded while trying to strip the anchor off.

    For example, if a page contains the link:

    /Company%20Info/103070.aspx

    It...

Monday, June 21 2010
28 100% cpu + memory filling at certain sites was updated
  • rb2k
    rb2k commented at 11:47 AM

    ok, just replaced nokogiri with hpricot, still the same problem

    This page is however weird: http://www.hiver2018.com/partenaires.php
    It's 34 MB and basically lo...

  • rb2k
    rb2k commented at 11:37 AM

    also: sites are the same and some sort of spam...
    they all lead to p*.vixns.eu

    my guess would be an xml parser going haywire

  • rb2k
    rb2k created the ticket at 11:33 AM

    I don't know what causes it yet, but there are some sites that seem to break the crawling process.
    two of them are e.g.
    http://www.skiandorre.com/
    http://www.hi...

Friday, June 11 2010
27 [patch] proxy was created
  • Rafał Lisowski
    Rafał Lisowski created the ticket at 6:10 AM

    Hi,
    Anemone do not support proxy. I create small patch to fix it.

Tuesday, June 01 2010
26 Illegal instruction was created
  • Robbie Clutton
    Robbie Clutton created the ticket at 12:33 PM

    Hi,

    I've just started using Anemone and I'm trying to use it in conjunction with a HTML/CSS validator. I'm getting the following error on my output when running...

Tuesday, May 25 2010
13 [patch] HTTP authentication was updated
15 Canonicalize URLs was updated
  • chris (at chriskite)
    • Milestone set to 0.5.0
    by chris (at chriskite) at 9:16 PM

    I think this makes sense if we utilize the rel=canonical tag on a page. The same script with a different query string should be considered as a different page u...

24 Memleak with pages? was updated
  • rb2k
    rb2k commented at 9:16 PM

    I crawl different domains though.
    Is there a way to "reset" the page cache in between crawls?

  • chris (at chriskite)
    • State changed from new to resolved
    by chris (at chriskite) at 9:13 PM

    Yes, although it's not really a "leak" because it is intentionally storing data about all the pages you crawl. If you crawl a lot of pages, that data has to go ...

21 pluggable html parser was updated
  • chris (at chriskite)
    • State changed from new to resolved
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 9:12 PM

    Thanks for your patch, I've incorporated your change to parse links with xpath instead of CSS.

    I don't plan to include pluggable parser support for a couple of ...

25 Normalize redirects was updated
22 option to stop crawling was updated
  • chris (at chriskite)
    • Milestone set to 0.5.0
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 9:06 PM
0.5.0 was created by chris (at chriskite) at 9:05 PM
 
3 of 3 tickets remain in this milestone
0.4.1 was created by chris (at chriskite) at 9:05 PM
 
0 of 1 tickets remain in this milestone
Tuesday, May 18 2010
25 Normalize redirects was created
  • rb2k
    rb2k created the ticket at 7:46 AM

    in the http.rb:

    redirect_to = response.is_a?(Net::HTTPRedirection) ?  URI(response['location']): nil
    
    should be
    redirect_to = response.is_a?(Net::HTTPRedirect...
Saturday, May 15 2010
24 Memleak with pages? was created
  • rb2k
    rb2k created the ticket at 4:07 PM

    Is it possible that, when running a lot of .crawl() operations, the @pages hash will keep on growing and there is no method that allows the user to empty it?
    (u...

Thursday, May 06 2010
23 Allow root domain redirects was created
  • rb2k
    rb2k created the ticket at 2:33 PM

    Anemone isn't currently able to crawl a site if the root domain redirects to another (sub)domain

    an example of this would be http://heise.de which redirects to ...

22 option to stop crawling was created
  • rb2k
    rb2k created the ticket at 10:56 AM

    It would be nice if there was a anemone.stop_crawl method that could be called from e.g. within the on_every_page block.
    In my case, I simply want to stop crawl...

Wednesday, May 05 2010
21 pluggable html parser was updated
  • rb2k
    rb2k commented at 12:33 PM

    Here are some performance benchmarks.
    Switching from CSS to xpath also results in speed boosts:

    http://gist.github.com/391134

    (tl;dr: hpricot (xpath) took: 3.59...

Tuesday, May 04 2010
21 pluggable html parser was created
  • rb2k
    rb2k created the ticket at 3:16 PM

    It would be nice if there was the possibility of using hpricot instead of nokogiri.
    They should be API compatible.

    In my tests, hpricot was always a tiny bit fa...

Thursday, February 04 2010
20 Crawl not staying on domain was updated
  • chris (at chriskite)
    • State changed from open to resolved
    by chris (at chriskite) at 11:23 PM

    Resolved in 0.3.2

  • chris (at chriskite)
    • State changed from new to open
    by chris (at chriskite) at 10:40 PM
  • Luke Hartman
    Luke Hartman created the ticket at 4:39 PM

    While running a crawl, I am sometimes getting page results from other domains. The following code:

    require 'rubygems'
    require 'anemone'
    require 'pp'
    
    Anemone.c...