Lighthouse has a new layout. Prefer the old one? Return to the old layout, and switch back any time from the link at the top of each page.

Activity (RSS)

Friday, January 22 2010
16 Page object doesn't expose @body was updated
13 [patch] HTTP authentication was updated
18 Support Multiple Encoding other than latin was updated
10 following <img> tag was updated
  • chris (at chriskite)
    • State changed from new to resolved
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 8:16 PM

    I plan to add an attr_accessor for Page body as part of ticket #16.

    To have Anemone visit img src's, use focus_crawl().

17 Anemone doesn't follow frame URLs was updated
  • chris (at chriskite)
    • State changed from new to open
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 8:12 PM
15 Canonicalize URLs was updated
  • chris (at chriskite)
    • State changed from new to open
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 8:12 PM
14 Error installing anemone was updated
  • chris (at chriskite)
    • State changed from new to resolved
    by chris (at chriskite) at 8:09 PM

    Indeed, this is an issue with installing nokogiri, which anemone depends on.

Wednesday, December 30 2009
19 How to save files matching regex? was created
  • dkelly
    dkelly created the ticket at 4:28 PM

    I'm new to Anemone and would like to ask how to get it to save all files who's URI matches a regular expression, eg. PDF files.

Monday, December 28 2009
18 Support Multiple Encoding other than latin was created
  • Kadvin
    Kadvin created the ticket at 1:49 AM

    The anemone should support multiple encoding, it involves:
    1. URL like: http://www.china.com/tag/中国 should be supported
    PS: URI(str_uri) will throw exception fo...

Saturday, December 26 2009
17 Anemone doesn't follow frame URLs was created
  • Noah Gibbs
    Noah Gibbs created the ticket at 5:28 PM

    I'm writing a simple spider to save bits of "http://api.rubyonrails.org". The top level page is a frameset, and links to a (nonexistent) HTML-only bit. It'd be ...

Thursday, December 24 2009
16 Page object doesn't expose @body was updated
  • Noah Gibbs
    Noah Gibbs commented at 3:34 PM

    Actually, it should probably be attr_reader rather than attr_accessor. Changing the body makes questionable sense in the first place, and doing it without chang...

  • Noah Gibbs
    Noah Gibbs created the ticket at 2:10 PM

    I'd like to use Anemone to mirror chunks of multiple web sites, and crawl to do it. I don't see any good way to get the actual page contents as text (rather tha...

Wednesday, December 16 2009
15 Canonicalize URLs was created
  • Nilesh
    Nilesh created the ticket at 11:14 PM

    It would be great if anemone can handle canonical URLs.

    For example, if we enable canonicalization,

      http://www.example.com/index.html
      http://www.example.co...
Friday, December 04 2009
14 Error installing anemone was created
  • mc
    mc created the ticket at 6:29 PM

    on a fresh ubuntu install 32bit

    gem install anemone

    Building native extensions. This could take a while...
    ERROR: Error installing anemone:

    ERROR: Failed to bu...
Friday, November 27 2009
13 [patch] HTTP authentication was updated
  • spk
    • Tag changed from authentication, http, patch to authentication, http
    by spk at 4:21 AM

    Hi,
    I have forked anemone from github, and add HTTP authentication, accessor for html body (needed for my project).
    http://github.com/spk/anemone/commit/420ff3...

Wednesday, November 18 2009
13 [patch] HTTP authentication was created
  • Bruno Michel
    Bruno Michel created the ticket at 12:12 PM

    Hi,

    Anemone ignores silently the username and password from a given URL. I've made a small patch to fix it.

Thursday, November 05 2009
11 Doesn't handle redirection properly. was updated
  • chris (at chriskite)
    • State changed from new to resolved
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 10:27 AM

    Since Anemone limits the crawl to a single domain, it won't switch over to your www subdomain after the redirect. You'll need to start on the domain you intend ...

Monday, November 02 2009
12 Thread(0x100190360): deadlock (fatal) was updated
  • chris (at chriskite)
    • State changed from new to resolved
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 11:47 AM

    0.2.3 includes a bugfix for this issue. Give it a try, feel free to reopen the ticket if you are still experiencing issues.

Friday, October 30 2009
12 Thread(0x100190360): deadlock (fatal) was updated
  • hayato
    hayato commented at 11:31 PM

    I have a similar experience of this ticket.

    Throwing exception from my code in result.

    Please add following flag to your code and retry.
    If error in your code, ...

  • Arun Thampi
    Arun Thampi created the ticket at 4:21 AM

    I get this error when I try to crawl URLs from a site. What can cause this? I'm running ruby 1.8.7 (2009-06-12 patchlevel 174) [i686-darwin10] on Mac OS X 10.6....

Saturday, October 24 2009
9 limit-depth crawling was updated
Monday, October 05 2009
11 Doesn't handle redirection properly. was created
  • urbanadventurer
    urbanadventurer created the ticket at 9:48 AM

    Doesn't handle redirection properly.

    irb(main):171:0 Anemone.crawl("http://treshna.com/") do |a|
    irb(main):172:1
    a.on_every_page do |x|
    irb(main):173:2* pp x
    ir...

Wednesday, September 16 2009
9 limit-depth crawling was updated
  • hayato
    hayato commented at 9:52 AM

    This ticket looks fixed at anemone 0.2.0.

    thanks!

10 following <img> tag was created
  • hayato
    hayato created the ticket at 9:51 AM

    I suggest adding feature to img tag tracking.

    • add :follow_img_tag option.
    • following img tag if :follow_img_tag => true.
    • add body attribute to Page to get raw h...
Sunday, September 06 2009
9 limit-depth crawling was updated
  • chris (at chriskite)
    • State changed from new to open
    • Assigned user set to chris (at chriskite)
    by chris (at chriskite) at 7:35 PM
  • hayato
    hayato created the ticket at 12:31 AM

    I suggest to add new function to anemone.

    Problem
    Anemone 0.1.2 follow link in same domain. In some of the cases of root URL, Anemone crawl too many pages, and ...

Monday, August 31 2009
4 github gem install doesn't include lib directory (broken) was updated
  • chris (at chriskite)
    • State changed from open to resolved
    by chris (at chriskite) at 10:28 PM

    Fixed the gemspec to work with GitHub. Decided not to distribute a gem through GitHub (using RubyForge only) to avoid confusion.

8 fail test by autotest was updated
  • chris (at chriskite)
    • State changed from new to resolved
    by chris (at chriskite) at 10:27 PM

    Thanks for the patch! I've applied this and pushed it to the GitHub repo.

Monday, August 24 2009
6 crawling slowly was updated
  • hayato
    hayato commented at 4:25 AM

    Your fix looks good.

    But, following test is failed.

    .F.FF.
    
    1)
    'Anemone::Page should store the response headers when fetching a page' FAILED
    expected nil? to r...
8 fail test by autotest was created
  • hayato
    hayato created the ticket at 4:04 AM

    Unittest of anemone 0.1.2 is written by RSpec.
    but test by autotest (included ZenTest) is failed.

    [~/git/anemone]> autospec
    loading autotest/rspec
    /opt/local/b...
Tuesday, August 18 2009
7 customizable page/link queue was created
  • hayato
    hayato created the ticket at 5:50 PM

    I suggest allowing customize queue.

    Current implementation page/link queue is instance of Queue object.
    But exist following use case.

    • use Message Queuing softw...
Monday, August 10 2009
6 crawling slowly was updated
  • chris (at chriskite)
    • State changed from new to resolved
    by chris (at chriskite) at 9:06 PM

    Hi,

    I've implemented time delay functionality in the latest release of Anemone (0.1.2). You can simply specify a :delay option when starting the crawl, like so:

    ...
Tuesday, August 04 2009
6 crawling slowly was updated
  • hayato
    hayato commented at 4:50 PM

    sorry my broken comment......

    Term "reserved" in patch causes confusion. "Delay" is felicity term.
    I attached the new patch.

    if apply this patch,this code behav...

  • hayato
    hayato commented at 4:41 PM

    Term "reserved" in patch causes confusion.
    "Delay" is felicity term.

    I attached the new patch.

    if apply this patch,this code behavior improve.

    require 'anemone'
    ...

Monday, August 03 2009
6 crawling slowly was updated
  • hayato
    hayato commented at 5:11 PM

    patch attatched

    add on_pre_fetch handler.

    if this block return false,link is not fetched and re-enq to link_queue.

  • hayato
    hayato created the ticket at 5:05 PM

    anemone is too violent to crawl any web server.

    example

    require 'anemone'
    Anemone.crawl("http://www.yahoo.co.jp/") do |anemone|
      anemone.on_every_page do |pag...
Wednesday, July 22 2009
5 Anemone doesn't use the query to test the page. was updated
Saturday, July 11 2009
4 github gem install doesn't include lib directory (broken) was updated
  • chris (at chriskite)
    • Assigned user set to chris (at chriskite)
    • State changed from new to open
    by chris (at chriskite) at 4:52 PM

    Current gemspec doesn't work with GitHub's SAFE=3 environment. Disabling GitHub gem support for the time being.

2 skip_links_if(&block) functionality was updated
  • chris (at chriskite)
    • State changed from new to resolved
    by chris (at chriskite) at 4:43 PM

    This functionality is covered by the new focus_crawl() method.

Friday, July 10 2009
4 github gem install doesn't include lib directory (broken) was created
  • Jeff
    Jeff created the ticket at 4:46 PM

    If you install via github you only get the bin directory. Without the lib directory, the gem doesn't work.

    Steps to reproduce:

    • sudo gem install chriskite-anemo...
Tuesday, July 07 2009
3 Won't crawl without a trailing forward slash on the url was created
  • Paul Barry
    Paul Barry created the ticket at 2:38 AM

    Something that tripped me up when I first tried to use anemone is to do:

    bin/anemone_url_list.rb http://paulbarry.com
    

    And I got back no results. I tracked it ...

Saturday, July 04 2009
3 Won't crawl without a trailing forward slash on the url was updated
Saturday, June 13 2009
1 OpenStruct functionality for Page objects was updated
Wednesday, June 10 2009
2 skip_links_if(&block) functionality was created
  • chris (at chriskite)
    chris (at chriskite) created the ticket at 4:59 PM

    It would be nice to be able to skip crawling pages based on more than just the URL matching a regex (skip_links_like). Something like this:

    anemone.skip_pages_...
1 OpenStruct functionality for Page objects was created
  • chris (at chriskite)
    chris (at chriskite) created the ticket at 4:56 PM

    It would be great to be able to do this:

    anemone.on_every_page do |page|
      page.h1 = (page.body).at('h1').inner_html
    end
    

    Probably this can be accomplished by ...

Home was created by chris (at chriskite) in Anemone at 4:40 PM