screen-scraper Quick Start Guide
Learn how to use screen-scraper in under 3 minutes.
Typical steps:


The best way to learn to use screen-scraper is by going through our tutorials.
These links allow you to get a general feel for screen-scraper. They are not representative of all that can be done. Each link will simple jump you to another section of this documentation.
screen-scraper will run on any operating system that supports version 1.5 or higher of the Java Virtual Machine. The installation process is almost always very simple, but you may want to read through these pages if you run into trouble or would like to know some related details.
The only specific restriction for installing screen-scraper is that your operating system platform supports a Java Runtime Environment of 1.8 or higher. Testing has been done for screen-scraper on Microsoft Windows, Linux, Mac OS X, and other platforms that support a Java Runtime Environment of 1.8 or higher. All systems have managed to run the software without any major changes.
The Windows installer comes with a runtime environment included. Mac OS X and Linux should already have a Java Runtime Environment installed. For help installing screen-scraper on other platforms (e.g., Solaris, FreeBSD) please contact us.
See also:
How much memory and what type of CPU is recommended for screen-scraper?
Scaling & Optimizing screen-scraper
To download screen-scraper, see the page for the edition you'd like to install. Run the installer to set installation options and install screen-scraper. For headless servers, run the installer with a -C flag to indicate the system is headless and the installer shouldn't try to open graphical popups for installation options
You may want to compare editions before choosing which to install.
screen-scraper ships with a Java Runtime Environment that should work in most distributions of Linux. Because the distributions can vary quite a bit, however, you may need to install a separate Java Runtime Environment and then point screen-scraper to it. If you're able to successfully install screen-scraper, but are having trouble starting it, try downloading the latest Java Runtime Environment from www.java.com for your particular distribution. Once you've installed the JRE, you can point screen-scraper to it by modifying the screen-scraper and server start scripts located in the screen-scraper installation folder.
As an example: the property might look like this:
This tells screen-scraper to use the JRE located at the given path rather than the one it ships with, or some other JRE that might be on your system.
screen-scraper License Agreement Copyright © 2002-2014 by ekiwi, LLC.
All Rights Reserved.
YOUR AGREEMENT TO THIS LICENSE
After reading this agreement carefully, if you ("Customer") do not agree to all of the terms of this End-User License Agreement ("EULA"), you may not use this Software (hereafter referred to as "Software Product"). Unless you have a different license agreement signed by ekiwi, LLC (hereafter referred to as "ekiwi") that covers this copy of the Software Product, your use of this Software Product indicates your acceptance of this EULA. All updates to the Software Product shall be considered part of the Software Product and subject to the terms of this EULA. Changes to this EULA may accompany updates to the Software Product, in which
case by installing such update Customer accepts the terms of the EULA as changed. The EULA is not otherwise subject to addition, amendment, modification, or exception unless in writing signed by an officer of both Customer and ekiwi. A software license and a license key ("Software Product License"), issued to a designated user only by ekiwi, is required for each concurrent user of the Software Product. By explicitly accepting this EULA you are acknowledging and agreeing to be bound by the following terms:
1. EVALUATION PERIOD
This Software Product may be used in conjunction with a free evaluation Software Product License. You may use the evaluation copy of the Software Product for only thirty (30) days in order to determine whether to purchase the Software Product, after which the Software Product will cease to function. ekiwi bears no liability for any damages resulting from use of the Software Product, and has no duty to provide any support before or after the expiration date of an evaluation license.
2. GRANT OF NON-EXCLUSIVE LICENSE
You may not tamper with, alter, or use the Software Product in a way that disables, circumvents, or otherwise defeats its built-in licensing verification and enforcement capabilities. You may not modify or create derivative copies of the Software Product or this EULA. All rights not expressly granted to you are retained by ekiwi.
ekiwi grants the non-exclusive, non-transferable right for a single user to use this Software Product. Each additional concurrent user of the Software Product must obtain an additional Software Product License. You may install the Software Product on as many computer systems as desired, so long as two copies of the same Software Product License never come into concurrent use.
3. INTELLECTUAL PROPERTY
The Software Product is owned by ekiwi and is protected by international copyright laws and treaties, as well as other intellectual property laws and treaties. You must not remove or alter any copyright notices on any copies of the Software Product. This Software Product copy is licensed, not sold. You may not use, copy, or distribute the Software Product, except as granted by this EULA, without written authorization from ekiwi. ekiwi reserves all intellectual property rights, including copyrights, patents, and trademarks.
4. TRANSFERABILITY
Customer may not rent, lease, lend, or in any way distribute or transfer any rights in this EULA or the Software Product to third parties without ekiwi's written approval, and subject to written agreement by the recipient of the terms of this EULA.
5. PROHIBITION ON REVERSE ENGINEERING AND DECOMPILATION
You may not reverse engineer, decompile, defeat license encryption mechanisms, or disassemble the Software Product or Software Product License except and only to the extent that such activity is expressly permitted by applicable law notwithstanding this limitation.
6. INDEMNIFICATION
You hereby agree to indemnify ekiwi against and hold harmless ekiwi from any claims, lawsuits, liabilty or other losses that arise out of your breach of any provision of this EULA.
7. THIRD PARTY SOFTWARE
Any software provided along with the Software Product that is associated with a separate license agreement is licensed to you under the terms of that license agreement (which license is provided with the Software Product). This license does not apply to those portions of the Software Product.
8. SUPPORT SERVICES
ekiwi may provide you with support services related to the Software Product. Use of any such support services is governed by ekiwi policies and programs described in online documentation and/or other ekiwi-provided materials.
As part of these support services, ekiwi may make available bug lists, planned feature lists, and other supplemental informational materials. ekiwi makes no warranty of any kind for these materials and assumes no liability whatsoever for damages resulting from any use of these materials. Furthermore, you may not use any materials provided in this way to support any claim made against ekiwi.
Any supplemental software code or related materials that ekiwi provides to you as part of the support services, in periodic updates to the Software Product or otherwise, is to be considered part of the Software Product and is subject to the terms and conditions of this EULA.
With respect to any technical information you provide to ekiwi as part of the support services, ekiwi may use such information for its business purposes without restriction, including for product support and development. ekiwi will not use such technical information in a form that personally identifies you without first obtaining your permission.
9. TERMINATION
This EULA terminates on the date of the first occurrence of either of the following events: (1) The expiration of one (1) month from written notice of termination from Customer to ekiwi; or (2) One party materially breaches any terms of this EULA or any terms of any other agreement between Customer and ekiwi, that are either uncorrectable or that the breaching party fails to correct within one (1) month after written notification by the other party.
10. NO WARRANTIES
YOU ACCEPT THE SOFTWARE PRODUCT AND SOFTWARE PRODUCT LICENSE "AS IS," AND EKIWI MAKES NO WARRANTY AS TO ITS USE, PERFORMANCE, OR OTHERWISE. TO THE MAXIMUM EXTENT PERMITTED BY APPLICABLE LAW, EKIWI DISCLAIMS ALL OTHER REPRESENTATIONS, WARRANTIES, AND CONDITIONS, EXPRESSED, IMPLIED, STATUTORY, OR OTHERWISE, INCLUDING, BUT NOT LIMITED TO, IMPLIED WARRANTIES OR CONDITIONS OF MERCHANTABILITY, SATISFACTORY QUALITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE, AND NON-INFRINGEMENT. THE ENTIRE RISK ARISING OUT OF USE OR PERFORMANCE OF THE SOFTWARE PRODUCT REMAINS WITH YOU.
11. LIMITATION OF CONSEQUENTIAL DAMAGES
NEITHER EKIWI NOR ANYONE INVOLVED IN THE CREATION, PRODUCTION, OR DELIVERY OF THIS SOFTWARE SHALL BE LIABLE FOR ANY INDIRECT, CONSEQUENTIAL, OR INCIDENTAL DAMAGES ARISING OUT OF THE USE OR INABILITY TO USE SUCH SOFTWARE EVEN IF EKIWI HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES OR CLAIMS. IN NO EVENT SHALL EKIWI'S LIABILITY FOR ANY DAMAGES EXCEED THE PRICE PAID FOR THE LICENSE TO USE THE SOFTWARE, REGARDLESS OF THE FORM OF CLAIM. EKIWI SHALL IN NO WAY BE HELD LIABLE OR RESPONSIBLE FOR ANY UNLAWFUL OR ILLEGAL USE OF THE SOFTWARE PRODUCT, INCLUDING, BUT NOT LIMITED TO, THE EXTRACTION AND USE OF COPYRIGHTED DATA FROM EXTERNAL SOURCES (E.G. WEB PAGES). THE PERSON USING THE SOFTWARE BEARS ALL RISK AND RESPONSIBILITY AS TO THE USE, QUALITY, AND PERFORMANCE OF THE SOFTWARE.
12. HIGH RISK ACTIVITIES
The Software Product is not fault-tolerant and is not designed, manufactured or intended for use or resale as on-line control equipment in hazardous environments requiring fail-safe performance, including, but not limited to, in the operation of nuclear facilities, aircraft navigation or communication systems, air traffic control, direct life support machines, and weapons systems, in which the failure of the Software Product, or any software, tool, process, or service that was developed using the Software Product, could lead directly to death, personal injury, or severe physical or environmental damage ("High Risk Activities"). Accordingly, ekiwi and its suppliers and licensors specifically disclaim any express or implied warranty of fitness for High Risk Activities. You agree that ekiwi and its suppliers and licensors will not be liable for any claims or damages arising from the use of the Software Product, or any software, tool, process, or service that was developed using the Software Product, in such applications.
13. GENERAL
This EULA is the complete statement of the agreement between the parties on the subject matter, and merges and supersedes all other or prior understandings, purchase orders, agreements and arrangements.
This EULA shall be governed by the laws of the State of Utah. Exclusive jurisdiction and venue for all matters relating to this EULA shall be in courts located in the State of Utah, and you consent to such jurisdiction and venue. If any action is brought by either party to this EULA against the other party regarding the subject matter hereof, the prevailing party shall be entitled to recover, in addition to any other relief granted, reasonable attorney fees and expenses of litigation.
You acknowledge that, in the event of your breach of any of the foregoing provisions, ekiwi will not have an adequate remedy in money or damages. ekiwi shall therefore be entitled to obtain an injunction against such breach from any court of competent jurisdiction immediately upon request. ekiwi's right to obtain injunctive relief shall not limit its right to seek further remedies.
There are no third party beneficiaries of any promises, obligations or representations made by ekiwi, LLC herein. Any waiver by ekiwi, LLC of any violation of this EULA by you shall not constitute or contribute to a waiver of any other or future violation by you of the same provision, or any other provision, of this EULA.
14. CONTACT INFORMATION
If you have any questions about this EULA, or if you want to contact ekiwi for any reason, please direct correspondence to [email protected].
screen-scraper's proxy server is a valuable tool for manipulating server interactions and building scrapes. In order for the proxy server to gather information from a browser, the browser must be configured to make requests to the server. As many people have never done this before we have provided instructions on how to setup some common browsers to use a proxy server. It is a simple process but one that might be a little foreign.
When you are done developing your sessions you will likely want to reset the proxy settings back to their normal state. This is not required so long as the proxy server is still running but once the proxy server is stopped the proxy settings will cause configured browsers to stop working until the settings are reset.
We've also added support to import proxy sessions from Charles proxy. You can proxy using Charles and then export the data to a "JSON Session File" and import that into screen-scraper. This can be a helpful alternative for proxying SSL sites.
Though any browser can be used to record transactions on the proxy server, we have found that some tend to experience less problems, complications, ans issues compared to others. On that note you might want to take some time to think about what browser you want to use when proxying a site.
If you are experiencing issues with transactions being recorded as errors, these can often be the results of browser plug-ins/add-ons. We have found that Internet Explorer is especially prone to them where as Opera tends to have the fewest issues.
Windows users:
.
If you have changed your proxy server settings to use a port other than 8777 then type your selected port in place of 8777.
If you're using a dial-up connection the setup will differ slightly. Instead of the LAN Settings button you'll want to find your dial-up connection under the Dial-up and Virtual Private Network settings dialog box, then configure it via the Settings button.
Depending on your operating system, instead of localhost you may need to use either 127.0.0.1 or the IP address of the machine. If you have trouble connecting to screen-scraper's proxy with your web browser, please see this FAQ.
Linux users:
.
If you're using a dial-up connection the setup will differ slightly. Instead of the LAN Settings button you'll want to find your dial-up connection under the Dial-up and Virtual Private Network settings dialog box, then configure it via the Settings button.
Depending on your operating system, instead of localhost you may need to use either 127.0.0.1 or the IP address of the machine. If you have trouble connecting to screen-scraper's proxy with your web browser, please see this FAQ.
Mac OS X users:
.
If you're using a dial-up connection the setup will differ slightly. Instead of the LAN Settings button you'll want to find your dial-up connection under the Dial-up and Virtual Private Network settings dialog box, then configure it via the Settings button.
Depending on your operating system, instead of localhost you may need to use either 127.0.0.1 or the IP address of the machine. If you have trouble connecting to screen-scraper's proxy with your web browser, please see this FAQ.
Windows users: click Options from the menu.
Linux users: click Preferences from the menu.
Mac OS X users: click Preferences from the menu.
If you have changed your proxy server settings to use a port other than 8777 then type your selected port in place of 8777.
If you're using a dial-up connection the setup will differ slightly. Instead of the LAN Settings button you'll want to find your dial-up connection under the Dial-up and Virtual Private Network settings dialog box, then configure it via the Settings button.
Depending on your operating system, instead of localhost you may need to use either 127.0.0.1 or the IP address of the machine. If you have trouble connecting to screen-scraper's proxy with your web browser, please see this FAQ.
For useful add-ons, visit the Browser Tools page.
If you have changed your proxy server settings to use a port other than 8777 then type your selected port in place of 8777.
If you're using a dial-up connection the setup will differ slightly. Instead of the LAN Settings button you'll want to find your dial-up connection under the Dial-up and Virtual Private Network settings dialog box, then configure it via the Settings button.
Depending on your operating system, instead of localhost you may need to use either 127.0.0.1 or the IP address of the machine. If you have trouble connecting to screen-scraper's proxy with your web browser, please see this FAQ.
Windows 7 users will need to select Preferences from the Settings menu after clicking to open Opera's menu.
Mac OS X users will need to select Preferences... from the Opera menu.
If you have changed your proxy server settings to use a port other than 8777 then type your selected port in place of 8777.
If you're using a dial-up connection the setup will differ slightly. Instead of the LAN Settings button you'll want to find your dial-up connection under the Dial-up and Virtual Private Network settings dialog box, then configure it via the Settings button.
Depending on your operating system, instead of localhost you may need to use either 127.0.0.1 or the IP address of the machine. If you have trouble connecting to screen-scraper's proxy with your web browser, please see this FAQ.
To simplify the whole process of turning the proxy server on and off, you can add a proxy button to a toolbar. One handy location is on the far right of the tab bar. To add the button,
You might need to select Appearance... explicitly before continuing.
New to Opera 9.5.x is a feature called Dragonfly. You can access it by selecting Developer Tools from > Advanced, or alternately pressing control-shift-I. If you're familiar with Firefox's Firebug add-on, then you'll quickly recognize Dragonfly. It's built off of the same ideas, allowing you to manipulate CSS in realtime, seeing which properties are being overwritten by another. You can debug Javascript running on the page, find HTML elements on the page by clicking on them so that Dragonfly shows you the corresponding page source, see all final properties of elements, etc. It's a great tool for sifting through a website.
Most of the general settings for screen-scraper are available through the workbench settings window. There are a handful of settings that are rarely used, so we don't provide a way to adjust them in the workbench. These properties can be edited manually in the screen-scraper.properties file in screen-scraper's resources/conf directory. You can edit it in your favorite text editor.
If running either the Basic Edition or Professional Edition note that when you alter the file you should do so when screen-scraper is not running. It won't get the new settings until the application restarts, and if you edit while it's running it may overwrite your changes.
If running the Enterprise Edition you have two options for reloading the screen-scraper.properties file while screen-scraper is running in server mode.
If you want to run multiple instances of screen-scraper on a single machine we advise that you modify and/or add several properties in the properties file (not just this one).
The screen-scraper.properties file can be found in the resource/conf directory of screen-scraper's installation. Most available settings and some sample values are listed below.
For the sake of readability the settings are listed in alphabetical order here. In a settings file they are ordered by a serialized value and so not alphabetical.

When running, the proxy server listens on a specified port for incoming HTTP requests from your web browser. Upon receiving a request from your browser the proxy server records it, then sends it along to the server for which it was intended. When that server responds it is received by the proxy server, which, once again, makes a record of it, then sends it along to your web browser.
screen-scraper's proxy server allows you to view HTTP requests and responses as they pass between your web browser and remote servers. In scraping files from web sites there are a few more details than you typically worry about when surfing, such as HTTP headers and POST data. The proxy server makes all of these details visible to you.
Often one of the headaches of scraping information from sites that use HTTPS is that it's not always easy to tell what's getting passed back and forth in the way of cookies, POST data, etc. Even if you put a proxy server in the way that lets you view the requests and responses, the information is encrypted as it's leaving your browser and as it's leaving the web server that responds to the request. screen-scraper gets around this problem by using it's own temporary certificate to encrypt traffic from itself to the browser and then encrypting each request before sending it up to the server. The result of this is that your browser will issue a warning about the certificate that screen-scraper returned. You can safely accept the certificate and be assured that all your traffic is encrypted.
We've also used other proxy software, such as Charles proxy, for handling SSL sites. They have additional features to allow the browser to trust the certificates so you don't see a warning in the browser. We've also added an import proxy session feature so you can import a JSON Session File from Charles and use it in screen-scraper to build a scraping session.
This feature is only available to Professional and Enterprise editions of screen-scraper.
screen-scraper has the ability to act as a proxy while in server mode. Combined with the ability to execute scripts, this functionality opens up many possibilities for how you use screen-scraper. More information about how to go about using screen-scraper in this capacity is available on our using scripts with the proxy server page.
First, create a proxy session to organize your interactions with the specific web sites.
Configuring a web browser to use a proxy server is generally pretty straightforward, but varies somewhat for each browser. We have provided instructions on how to setup different browsers:
Assuming you've configured everything and set up a proxy session, from here you should be able to start up the proxy server by selecting your proxy session in the objects tree and then clicking on the Start Proxy Server button in the general tab. Now just surf the pages that you want to record.
After you've surfed a bit with your web browser click on the progress tab. From here you can view all of the HTTP and HTTPS requests and responses logged by the proxy server. Clicking on a transaction brings up its details in the lower pane.
If you are using Internet Explorer 7 you have to adjust your security settings. To do this open Internet Options in the menu and under the security tab change the security level to medium.
If security settings are not updated you will see an error page when accessing a site that uses HTTPS encryption.

IE domain mismatch warning
This warning occurs because screen-scraper is using a temporary certificate for encryption that will not match the url that you are accessing. You can safely ignore this warning by clicking Continue to this website. This practice is, however, not recommended.
Most browsers have recently started preventing you from accepting the certificate. As a work-around you can use another proxy service, such as Charles, to proxy SSL sites by installing their root authority certificate on your server (which prevents an error from showing up while using them). You can then import the proxy data into screen-scraper for building your scrape
If you normally use an external proxy server when connecting to the internet (on your local area network, for example), you'll need to specify this information in screen-scraper's external proxy settings. Before you can run the proxy server.
This feature has been deprecated and by default is not available in the workbench interface. To enable proxy scripting please add AllowProxyScripting=true to your resource/conf/screen-scraper.properties file and restart screen-scraper.
screen-scraper has the ability to run custom-made scripts while the proxy server is running (more information on starting and stopping the server is available). This allows you to setup blacklists, filter web pages, or otherwise manipulate browser requests and server responses. It is recommended that you read about managing and using scripts before continuing.
The scripts tab is used to associate scripts with a proxy. Depending on when you decide to run your script, certain built in objects will be in scope that are unique to the proxy environment.
screen-scraper offers a few objects that you can work with in a script in the proxy environment. See the variable scope section and/or API documentation for more details.
Depending on when a script gets run different variables may be in scope (available). The table that follows specifies what variables will be in scope depending on when a given script is run.
| When Script is Run | proxySession in scope | request in scope | response in scope |
|---|---|---|---|
| Beginning of proxy session | X | ||
| Before HTTP request | X | X | |
| After HTTP request | X | X | |
| Before HTTP response | X | X | X |
| After HTTP response | X | X | X |
One of the best ways to fix errors is to simply watch the proxy session log (under the log tab in the proxy session) and the error.log file (located in the log directory of screen-scraper's install directory) for script errors. When a problem arises in executing a script screen-scraper will output a series of error-related statements to the logs. Often a good approach in debugging is to build your script bit by bit, running it frequently to ensure that it runs without errors as you add each piece.
This feature is only available to Professional and Enterprise editions of screen-scraper.
It is strongly advised NOT to run both the server and workbench simultaneously.
There are two main reasons to run screen-scraper as a server:
If you're running Microsoft Windows screen-scraper will run as a service. This allows it to be run in a background process, and doesn't require you to be logged in to the machine on which it's running. screen-scraper gets registered as a service upon installation, and may be run as a server using either the Start server and Stop server links from the menu, or using the Services control panel applet.
In Windows XP, when the server is running an icon will appear in the system tray. You can right-click this icon to stop the server.
The screen-scraper service can be started, stopped, and monitored via the Services control panel applet under Administrative Tools. The server can also be started and stopped via the Start server and Stop server shortcuts found under the menu.
When the server is running, the system tray icon will not appear, as it does in other versions of Windows.
There are a few ways to determine whether or not the server is running in Vista:
As of 5.0 there are three batch files that have been added to screen-scraper to aid in managing the server from the command line.
These files are created when screen-scraper installs and keyed specifically to the new instance of screen-scraper.
Under Unix/Linux or Mac OS X the server is controlled via the server script, which operates much like a typical Unix daemon. The server will run in a background process, allowing you to start and stop the server remotely, or log out of your session after starting it. You can issue the following commands to the server script:
While screen-scraper is running as a server it will accept connections from any other machine unless you specify otherwise. It is important to consider the security of your machine when running any service of this type. screen-scraper allows you to specify the IP addresses of the machine(s) you wish to allow to connect to it in the IP addresses to allow to connect in the settings window.
In this field it expects a comma-delimited list of IP addresses that screen-scraper should accept connections from. You can also specify just the beginning portions of IP addresses. For example, if you enter 111.22.333 screen-scraper would accept connections from 111.22.333.1, 111.22.333.2, 111.22.333.3, etc.
If nothing is entered into this text box screen-scraper will accept connections from any IP address. This is discouraged unless the computer running screen-scraper is protected by an external firewall.
If you need to alter this setting in a GUI-less environment, you can close screen-scraper and edit the resource/conf/screen-scraper.properties file. The setting to change is IPAddressesToAllow. When you start screen-scraper, it will make use of the new setting.
If you're having trouble starting screen-scraper in server mode or running scraping sessions in server mode please see our FAQ on troubleshooting server mode issues.
The scripting engine requests files which it then parses, manipulates, and/or stores according to user defined processes. It is the heart of screen-scraper and has been optimized at all points of development to be as efficient as possible. It is made up of multiple parts which can be manipulated using the workbench:
The rest of this section contains information about using screen-scraper, through the workbench, to achieve different goals. These can be difficult to understand without some exposure to the software. That is why we would like to encourage you to go through our first few tutorials before continuing.
screen-scraper allows for Java libraries to be added to the classpath. Simply copy any jar files you'd like to reference into the lib\ext folder found in screen-scraper's directory. The next time you start up screen-scraper it will automatically add the jar files to its classpath. Note that you'll still need to use the import statement within your scripts to refer to specific classes:
screen-scraper was built on a Java 1.5 platform. Your Java scripts must accept at least a version 1.5 JRE in order to compile and run properly.
Under certain circumstances you may want to anonymize your scraping so that the target site is unable to trace back your IP address. For example, this might be desirable if you're scraping a competitor's site, or if the web site blocks IP addresses that make too many requests.
There are a few different ways to go about this using screen-scraper:
If you choose to run anonymous scripts from an external script, it is valuable to read through the documentation on controlling anonymization externally.
Aside from the above methods, you might find our blog posting on how to surf and screen-scrape anonymously helpful. It's slightly dated, but still very relevant.
The screen-scraper automatic anonymization service works by sending each HTTP request made in a scraping session through a separate high-speed HTTP proxy server. The end effect of this is that the site you're scraping will see any request you make as coming from one of several different IP addresses, rather than your actual IP address. These HTTP proxy servers are actually virtual machines that get spawned and terminated as you need them. You'll use screen-scraper to either manually or automatically spawn and terminate the proxy servers.
Note: When using the automatic anonymization method, while the remote web site may not be able to determine your IP address, your activity will still be logged. If you attempt to use the proxy service for any illegal activities, the chances are very good that you will be prosecuted.
While the automatic anonymization service provides an excellent way to cloak your IP address it is still possible that the target web site will block enough of the anonymized IP addresses that the anonymization could fail. Unfortunately we can't make any guarantees that you won't get blocked; however, by using the automatic anonymization service the chances of getting blocked are reduced dramatically.
The anonymous proxy servers will be set up in such a way that they only allow connections from your IP address. This way no one else can use any of the proxies without your authorization. This configuration is tied to your password. For more on restricting connections see documentation on managing the screen-scraper server.
If you'll be running your anonymized scraping sessions on the same machine (or local network) you're currently on and you are using the workbench, you can click the Get the IP address for this computer button to determine your current IP address.
Anonymization settings can be configured using screen-scraper's workbench. Settings are determined in the anonymous proxy settings of the settings dialog box.

When you sign up for the anonymization service you'll be given the password that allows your instance of screen-scraper to manage anonymous proxies for you. You'll enter it into the Password textbox in the settings.
As the proxy servers get spawned and terminated, it's a good idea to establish the maximum number of running proxy servers you'd like to allow. This is done via the Max running servers setting. Because you pay for proxy servers by the hour, if you don't have your scraping session set up to automatically shut them down at the end, you'll use the Terminate all running proxy servers button in order to do that.
We find that as many as 10 proxy servers but no fewer than five are adequate for most situations.
If you're setting this value in a GUI-less environment (i.e., a server with no graphical interface), you'll want to set these values in the resource/conf/screen-scraper.properties file (if these property is not already in the file you'll want to add it).
Acceptable values are http://anon.screen-scraper.com and http://anon2.screen-scraper.com.
Be sure to modify the resource/conf/screen-scraper.properties file only when screen-scraper is not running.
Aside from these global settings, there are a few settings that apply to each scraping session you'd like to anonymize. You can edit these settings under the anoymization tab of your scraping session.
Once you've configured all of the necessary settings, try running your scraping session to test it out. You'll see messages in the log that indicate what proxy servers are being used, how many have been spawned, etc.
As your anonymous scraping session runs, you'll notice that screen-scraper will automatically regulate the pool of proxy servers. For example, if screen-scraper gets a timed out connection or a 403 response (authorization denied), it will terminate the current proxy server, and automatically spawn a new one in its place. This way you will likely always have a complete set of proxy servers, regardless of how frequently the target web site might be blocking your requests. You can also manually report a proxy server as blocked by calling session.currentProxyServerIsBad() in a script. When this method is called the current proxy server will be shut down and replaced by another.
If the automatic anonymization method isn't right for you, the next best alternative might be to manually handle working with screen-scraper's built-in ProxyServerPool object. The basic approach involves running a script at the beginning of your scraping session that sets up the pool, then calling session.currentProxyServerIsBad() as you find that proxy servers are getting blocked. In order to use a proxy pool, you'll also need to get a list of anonymous proxy servers. Generally you can find these by Googling around a bit.
See available methods:
ProxyServerPool
Anonymization API
That's about all there is to it. Aside from occasionally calling session.currentProxyServerIsBad(), you may also want to call session.setUseProxyFromPool to turn anonymization on and off within the scraping sesison.
The web interface is only available for enterprise edition users of screen-scraper.
The mapping tab allows you to alter extracted values. Often once you extract data from a web page you need to put it into a consistent format. For example, you may want products with very similar names to have identical names.
screen-scraper makes use of mapping sets when determining how to map a given extracted value. A mapping set may contain any number of mappings, which screen-scraper will analyze in sequence until it finds a match, or runs out of mappings. As such, you'll often want to put more specific mappings higher in sequence than more general mappings.

Consider the screen-shot of the mapping tab: if the extracted value were Widget 123 screen-scraper would first try to match using the Widget 1 mapping. Because this is an equals match the mapping wouldn't occur, so screen-scraper would proceed to the second mapping. The second mapping would match because a contains type was designated. That is, the text Widget 123 contains the text Widget. As such, the extracted data Widget 123 would become Product ABC, because that is the To value designated for the second mapping.
When using regular expressions in your mapping you can also make use of back references. Back references allow you to preserve values in the original text when mapped to the To value. For example, if you were mapping the value Widget 123 you could use the regular expression Widget (\d*). In the To column you could then enter the value Product \1, which, when mapped, would convert Widget 123 to Product 123. The value in parentheses in the From column gets inserted via the \1 marker found in the To column.
This feature is only available to Professional and Enterprise editions of screen-scraper.
It is possible to run a scraping session within a scraping session that is already running. This is done with the RunnableScrapingSession class. Detailed documentation on methods available for the RunnableScrapingSession class are available in our API documentation. Here's a specific example of how the RunnableScrapingSession might be used in a screen-scraper script:
Scripts attached to a scraping session are exported along with it. When you subsequently import that scraping session into another instance of screen-scraper it might overwrite existing scripts in that instance. In some cases, though, you might have a series of general scripts shared by many scraping sessions. In these cases you often want to ensure that the very latest versions of these general scripts get retained in a given instance.
In the main pane of the script there is a Overwrite this script on import checkbox. When checked, any name clashes between existing scripts and imported versions will prompt you whether it should overwrite the script or not. If the local version's Overwrite this script on import checkbox is unchecked this file will not be overwritten even if you click to have it overwritten.
Checking the Overwrite this script on import requires access to the screen-scraper workbench, which you may not have access to if screen-scraper is running in a GUI-less environment. In these cases you can make use of the ForceOverwriteScripts property in the resource/conf/screen-scraper.properties file to allow scripts that have this box un-checked to be overwritten. In order to overwrite scripts that have this checkbox un-checked in a GUI-less environment you would follow these steps:
Once you're finished importing you may want to stop screen-scraper, set the property back to false, then start up again. Note that when you import scripts with the ForceOverwriteScripts property set to true screen-scraper will import the scripts regardless of whether or not the Overwrite this script on import checkbox is checked.
Extractor patterns allow you to pinpoint select snippets of data that you want extracted from a web page. It is a block of text (usually HTML) that contains special tokens that will match pieces of data you're interested in extracting. These tokens are text labels surrounded by the delimiters ~@ and @~ (e.g. ~@NAME@~). The identifier between the delimiters can contain only alpha-numeric characters and underscores.
Extractor patterns are added to scrapeable files under the extractor patterns tab.
You can think of an extractor pattern like a stencil. A stencil is an image in cut-out form, often made of thin cardboard. As you place a stencil over a piece of paper, apply paint to it, then remove the stencil, the paint remains only where there were holes in the stencil. Analogously, you can think of placing an extractor pattern over the HTML of a web page. The tokens correspond to the holes where the paint would pass through. After an extractor pattern is applied it reveals only the portions of the web page you'd like to extract.
Extractor tokens designate regions where data elements are to be captured. For example, given the following HTML snippet:
you would extract piece of text by creating an extractor pattern with a token positioned like so:
The extracted text could then be accessed via the identifier EXTRACTED_TEXT.
If you haven't done so already, we'd recommend going through our first tutorial to get a better feel for using extractor patterns.
If an extractor pattern takes too long to match a block of text it will timeout. The timeout setting may be adjusted from the general tab of the Settings located in the menu. If you find that your extractor pattern is timing out you might try adjusting it by using more precise regular expressions.
screen-scraper's scraping engine allows you to associate custom scripts with various events in the scraping process. It is recommended that you read about managing and scripting in screen-scraper before continuing.
Depending on what event triggers a script to be run different objects will be in-scope. Triggers regarding the scraping session are added on the general tab of the scraping session, file request/response triggers are associated on the properties tab of the scrapeable file, and extractor pattern events in the scripts section of the main tab in the extractor patterns tab of the scrapeable file.
Scripts can also be used to run scripts using the session.executeScript method.
screen-scraper offers a few objects that you can work with in a script in the scraping engine. See the variable scope section and/or API documentation for more details.
Depending on when a script gets run different variables may be in or out of scope. When associating a script with an object, such as a scraping session or scrapeable file, you're asked to specify when the script is to be run. The table that follows specifies what variables will be in scope depending on when a given script is run. Only variables that are in scope are accessible to the script.
| When Script is Run | session in scope | scrapeableFile in scope | dataSet in scope | dataRecord in scope |
| Before scraping session begins | X | |||
| After scraping session ends | X | |||
| Before file is scraped | X | X | ||
| After file is scraped | X | X | ||
| Before pattern is applied | X | X | ||
| After pattern is applied | X | X | X | |
| Once if pattern matches | X | X | X | X |
| Once if no matches | X | X | ||
| After each pattern match | X | X | X | X |
One of the best ways to fix errors is to simply watch the scraping session log and the error.log file (located in the log directory where screen-scraper was installed) for script errors. When a problem arises in executing a script screen-scraper will output a series of error-related statements to the logs. Often a good approach in debugging is to build your script bit by bit, running it frequently to ensure that it runs without errors as you add each piece.
When screen-scraper is running as a server it will automatically generate individual log files in the log directory for each running scraping session (this can be disabled in the settings window). An error.log file will also be generated in that same directory when internal screen-scraper errors occur.
The breakpoint window can also be invaluable in debugging scripts. You can invoke it by inserting the line session.breakpoint() into your script.
Session variables allow you store values that will persist across the life of a scraping session.
There are a few different ways to set session variables.
As with setting session variables, there is more than one way to retrieve values of session variables.
If you have a session variable identified by QUERY_PARAM you might embed it into the URL field of a scrapeable file using http://www.mydomain.com/myscript.php?query=~#QUERY_PARAM#~. screen-scraper will automatically replace the ~#QUERY_PARAM#~ with the value of the session variable.
Sub-extractor patterns allow you to extract data in the context of an extractor pattern, providing significantly more flexibility in pinpointing the specific pieces you're after. Consider a search results page consisting of rows and columns of data. Using normal extractor patterns you would use a single pattern to extract the data from all columns for a single row. In many cases this works just fine; however, the process gets more complicated when each row differs significantly. For example, certain cell rows may be in different colors or their contents may be completely missing. With a normal extractor pattern it would be difficult to account for the variability in the cells. By using sub-extractor patterns you could create a normal extractor pattern to extract an entire row, then use individual sub-extractor patterns to pull out the individual cells.
When using sub-extractor patterns only the first match will be used. That is, even if a sub-extractor pattern could match multiple times, only the data corresponding to the first match will be extracted. Because of this sub-extractor patterns are not always the correct method for getting data within a larger context. To get multiple matches in a larger context, like all rows in a table, you would instead use manual extractor patterns.
Consider the following HTML table:
| Name | Phone | Address |
|---|---|---|
| Juan Ferrero | 111-222-3333 | 123 Elm St. |
| Joe Bloggs | No contact information available | |
| Sherry Lloyd | 234-5678 (needs area code) | 456 Maple Rd. |
Here is the corresponding HTML source:
It would be difficult to write a single extractor pattern that would extract the information for each row because the contents of the cells differ so significantly. The different colored cells and the cell spanning two columns make the data too inconsistent to be easily extracted using a single pattern (which would require lots of regular expressions and might still prove impossible or inconsistent).
Consider this extractor pattern:
The ~@DATARECORD@~ extractor pattern token is special in that it defines the block of data to which you wish to apply sub-extractor patterns. Sub-extractor patterns cannot be applied to a token with a name other than DATARECORD
If applied to the HTML above the extractor pattern would produce the following three matches:
Sub-extractor patterns would allow you to extract individual pieces of information from each row. For example, consider this sub-extractor pattern:
If applied to each of the individual extracted rows above the following three pieces of information would be extracted:
This is a simple case. Now consider the extractor pattern for the phone number:
If applied to each of the individual extracted rows above the following three pieces of information would be extracted:
In the case of Sherry Lloyd this presents a serious problem because she does have a phone number listed. It is not selected because of the additional class. Let's adjust the sub-extractor pattern slightly:
The ~@nondoublequotes@~ represents an extractor token that uses the Non-double quotes regular expression: [^"]*. Matching anything between where it is covering until it encounters double quotes. In this particular case Sherry's phone number also gets extracted.
We now have the case of the cell in the second row that spans two columns, which would not get extracted by our current sub-extractor patterns. We may still want this information, however, so we create the following sub-extractor pattern, just in case the cell exists:
If applied to our data we'd get the following results:
When multiple sub-extractor patterns hold a token with the same name (in this case, PHONE), the last one to match is the one that determines the value of the token. In this example either one or the other will match. If both could match then we would want to have the first phone extractor pattern ordered later than the one to match the no-data-available pattern
Sub-extractor patterns aggregate everything that's extracted into a single data set. Using all of our extractor and sub-extractor patterns together we'd get the following data set:
| Data record # | Name | Phone |
|---|---|---|
| Data record #1 | Juan Ferrero | 111-222-3333 |
| Data record #2 | Joe Bloggs | No contact information available |
| Data record #3 | Sherry Lloyd | 234-5678 (needs area code) |
To launch the Workbench, double-click the screen-scraper icon
The workbench provides an intuitive and convenient way to interact with screen-scraper's scraping engine. This section of our documentation covers the interfaces provided in the workbench to develop and manage scrapes. If you're interested in learning to use screen-scraper, the best approach is to go through at least our first few tutorials.
To ensure clarity, the first thing that you need to know about the workbench is the names for the various regions of the window. This will help to keep us oriented correctly during the documentation of the workbench.
If this is your first time opening screen-scraper the only item listed will be the root folder.

The size of the two panes can be adjusted to your liking by clicking on the vertical bar that divides the two panes and dragging it to the left or right.
This section contains a description of each of the screens found in the Settings window, which can be displayed by selecting Settings from the menu, or by clicking the wrench icon in the button bar.


These settings apply when screen-scraper is running in server mode.
These settings apply only to the proxy server portion of screen-scraper.
These settings are used with the sutil.sendMail method in screen-scraper scripts.
These settings apply only to the web interface and SOAP server features of screen-scraper.

Unless you normally connect to the Internet through an external proxy server, you don't need to modify these settings.
If you are using NTLM (Windows NT) authentication you'll need to designate settings for both the standard proxy as well as the NTLM one.

This setting is available in the screen-scraper.properties file as AnonymousProxyPassword
In this field it expects a comma-delimited list of IP addresses that screen-scraper should accept connections from. You can also specify just the beginning portions of IP addresses. For example, if you enter 111.22.333 screen-scraper would accept connections from 111.22.333.1, 111.22.333.2, 111.22.333.3, etc.
If nothing is entered into this text box screen-scraper will accept connections from any IP address. This is not generally encouraged.
This setting is available in the screen-scraper.properties file as AnonymousProxyAllowedIPs
This setting is available in the screen-scraper.properties file as AnonymousProxyMaxRunning
As you pay for proxy servers by the hour, if you don't have your scraping session set up to automatically shut them down at the end you will need to use this button to end the proxy servers.
Under certain circumstances you may want to anonymize your scraping so that the target site is unable to trace back your IP address. For example, this might be desirable if you're scraping a competitor's site, or if the web site is blocking too many requests from a given IP address.
There are a few different ways to go about this using screen-scraper. We will discuss how to setup anonymazation in screen-scraper later in the documentation.
A proxy session in screen-scraper is a record of the requests and responses that go between a browser and a proxy server. It is useful in learning how to scrape a site and is used to configure screen-scraper's proxy server. For more information see our documentation about using the proxy server.


For the button to work correctly you will want to clear your browser cookies before having the proxy session record all transactions. This makes it so that cookies already in existence are not considered to be javascript cookies.
Transactions not included in the list are still recorded to the proxy session log.
When a transaction is selected more information regarding the request and response is displayed.


screen-scraper has always kept track of server set cookies and does that for you automatically; however, when the cookies are set by javascript screen-scraper does not catch them. This saves on the time lost having screen-scraper scrape every javascript file when most of the time there is nothing there that matters.
This mean that you have to set any javascript added cookies using the setCookie method. To help find where javascript cookies are being set we have added a button in the proxy session progress tab.
For the button to work correctly you will want to clear your browser cookies before having the proxy session record all transactions. This makes it so that cookies already in existence are not considered to be javascript cookies.
This feature has been deprecated and by default is not available in the workbench interface. To enable proxy scripting please add AllowProxyScripting=true to your resource/conf/screen-scraper.properties file and restart screen-scraper.
You are unlikely to use this tab unless you are running screen-scraper as a proxy in server mode.


If you are trying to troubleshoot problems with scripts not working the way you expected the log can give you clues as to where problems might exists. Likewise, you can have your scripts write to the log to help identify what they are doing. If you have selected to filter out binary files and/or less useful transactions a log of those transactions will be available here.
The proxy session log is not saved in the workbench, if you close screen-scraper you will lose the current contents of the proxy session log.

Some web sites require that you supply a client certificate, that you would have previously been given, in order to access them. This feature allows you to access this type of site while using screen-scraper.
For more info see our blog entry on the topic.
A scraping session is simply a way to collect together files that you want scraped. Typically you'll create a scraping session for each site from which you want to scrape information.
screen-scraper should not be running when you add the file into the folder. All files will be imported into the root folder the next time screen-scraper starts.
When a scraping session is exported it will use the character set indicated under the advanced tab. If a value isn't indicated there it will use the character set indicated in the general settings.

Each script can be designated to run either before or after the scraping session runs. This can be useful for functions like initializing session variables and performing clean-up after the scraping session is finished. It's often helpful to create debugging scripts in your scraping session, then disable them once you're ready to run your scraping session in a production environment.

If you are trying to troubleshoot problems with scripts not working the way you expected the log can give you clues as to where problems might exists. Likewise, you can have your scripts write to the log to help identify what they are doing.
This tab displays messages as the scraping session is running. This is one of the most valuable tools in working with and debugging scraping sessions. As you're creating your scraping session you'll want to run it frequently and check the log to ensure that it's doing what you expect it to.

There may be instances where you find yourself unable to log in to a web site or advance through pages as you're expecting. If you've checked other settings, such as POST and GET parameters, you may need to adjust the cookie policy. Some web sites issue cookies in uncommon ways, and adjusting this setting will allow screen-scraper to work correctly with them.
If pages are rendering with strange characters then you likely have the wrong character set. You should also try turning off tidying if international characters aren't being rendered properly.
Some web sites require that you supply a client certificate, that you would have previously been given, in order to access them. This feature allows you to access this type of site while using screen-scraper.
If you are using NTLM (Windows NT) authentication you'll need to designate settings for both the standard proxy as well as the NTLM one.

Should proxy servers fail to spawn screen-scraper will proceed forward with the scraping session once at least 80% of the minimum required proxy servers are available.
This tab is specific for automatic anonymization For more information on anonymization, see our page on how to set it up anonymization in screen-scraper.
A scrapeable file is a URL-accessible file that you want to have retrieved as part of a scraping session. These files are the core of screen-scraping as they determine what files will be available to extract data from.
In addition to working with files on remote servers, screen-scraper can also handle files on local file systems. For example, the following is a valid path to designate in the URL field: C:\wwwroot\myweb\my_file.htm.

You can tell what files are being scraped manually and which are in sequence using the objects tree. Sequenced scrapeable files are displayed with a pound sign (#) on them.

GET parameters can also be embedded in the URL field under the Properties tab.
Parameters can be deleted by selecting them and either hitting the Delete key on the keyboard, or by right-clicking and selecting Delete.
Session variables can be used in the Key and Value fields. For example, if you have a POST parameter, username, you might embed a USERNAME session variable in the Value field with the token ~#USERNAME#~. This would cause the value of the USERNAME session variable to be substituted in place of the token at run time.
In the enterprise edition of screen-scraper you can also designate files to be uploaded. This is done by designating FILE as the parameter type. The Key column would contain the name of the parameter (as found in the corresponding HTML form), and the value would be the local path to the file you'd like to upload (e.g., C:\myfiles\this_file.txt).

This button is grayed out if there is not a extractor pattern currently copied.
This tab holds the various extractor patterns that will be applied to the HTML of this scrapeable file. The inner frame will be discussed in more detail when discussing them.

This can be very helpful for pages that are very specific on request settings or where you are getting unexpected results from the page. This is the best place to start when you experience this type of issue.
This tab will display the raw HTTP request for the last time this file was retrieved. This tab can be useful for debugging and looking at POST and GET parameters that were sent to the server.

The contents shown under the this tab might appear differently from the original HTML of the page. screen-scraper has the
ability to tidy the HTML, which is done to facilitate data extraction. See using extractor patterns for more details.
The most common use for this tab is in generating and testing extractor patterns. You can generate
an extractor patterns by highlighting a block of text or HTML, right-clicking and selecting
Generate extractor pattern from selected text.

You can generally recognize when a web site requires this type of authentication because, after requesting the page, a small box will pop up requesting a username and password.
A minor performance hit is incurred, however, when tidying. In cases where performance is critical Don't Tidy HTML should be selected.
Extractor patterns allow you to pinpoint snippets of data that you want extracted from a web page. They are made up of text (usually HTML), extractor tokens, and possibly even session variables. The text and session variables give context to the tokens that represent the data that you want to extract from the page.
Extractor patterns can be difficult to understand at first. We recommend that you read about using extractor patterns or go through our first tutorial before continuing.
When creating extractor patterns you should use the HTML that will be found under the last response tab associated with a scrapeable file. By default, screen-scraper will tidy the HTML once it's been scraped, meaning that it will format it in a consistent way that makes it easier to work with. If you use the HTML by viewing the source for a page in your web browser it will likely be different from the HTML that screen-scraper generates.


The buttons specific to the sub-extractor pattern are discussed in more detail later in this documentation.

It is recommend that you generally avoid checking this box unless it's absolutely needed because of memory issues it may cause. If this box is checked, screen-scraper will continue to append data to the dataSet, and all of that data will be kept in memory. The preferred method is to save data as it's being extracted, generally by invoking a script with a script association After each pattern match that pulls the data from dataRecord objects or session variables.
Sub-extractor patterns allow you to extract data within the context of an extractor pattern, providing significantly more flexibility in pinpointing the specific pieces you're after. Please read our documentation on using sub-extractor patterns before deciding to use them.
Sub-extractor patterns only match the first element they can. To get multiple matches, you would use manual extractor patterns instead.

Extractor tokens select the information from a file that you want to be able to access. The purpose of an extractor pattern is to give context to the extractor token(s) that it contains. This is to assist in getting the tokens to only return the information that you desire to have. Without extractor tokens you will not gather any information from the site.
Extractor tokens become available to dataRecord, dataSet, and session objects depending on their settings and the scope of the scripts invoked. All extractor tokens are surrounded by the delimiters ~@ and @~ (one for each side of the token). Between the two delimiters is where the name/identifier of the token is specified.
Make appropriate changes to the TOKEN_NAME text to reflect the desired name of the token.

The regular expressions that appear in the drop-down list can be edited by selecting Edit regular expressions from the menu.

We would encourage you to read our documentation on mapping extracted data before you start using mappings.
To create a new set, select the text in the Set textbox and start typing the name of the new set.
Mappings can be deleted by pressing the Delete key on your keyboard after selecting them.

screen-scraper has a built-in scripting engine to facilitate dynamically scraping sites and working with data once it's been extracted. Scripts can be helpful for such things as interacting with databases and dynamically determining which files get scraped at when.
Invoking scripts in screen-scraper is similar to other programming languages in that they're tied to events. Just as you might designate a block of code to be run when a button is clicked in Visual Basic, in screen-scraper you might run a script after an HTML file has been downloaded or data has been extracted from a page. For more information see our documentation on scripting triggers.
Depending on your preferences, there are a number of languages that scripts can be written in. You can learn more in the scripting in screen-scraper section of the documentation.
If you haven't done so already, we'd highly recommend taking some time to go through our tutorials in order to get more familiar with how scripts are used.
If screen-scraper is running when you copy the files into the import folder they will be imported and hot-swapped in the next time a scraping session is invoked. They will also be imported if you start or stop screen-scraper.

For example, scripts attached to a scraping session are exported along with it. When you subsequently import that scraping session into another instance of screen-scraper it might overwrite existing scripts in that instance. For more information read our documentation on script overwriting.
You designate a script to be executed by associating it with some event. For example, if you click on a scraping session, you'll notice that you can designate scripts to be invoked either before a scraping session begins or after it completes. Other events that can be used to invoke scripts relate to scrapeable files and extractor patterns.

Available associations (based on object location) are listed with a brief description of how they can be useful.
All objects that can have scripts associated with them have buttons to add the script association with the exception of scripts. To create a association between scripts you would use the executeScript method of the session object.
Locations to specify script associations are listed below.
Script associations are ordered automatically in a natural order based on their relation to the object they are connected to: scripts called after the file is scraped cannot be ordered before associations the are called before the file is scraped. Beyond the natural ordering you can specify the order of the scripts using the Sequence number.
You can selectively enable and disable scripts using the Enabled checkbox in the rightmost column. It's often a good practice to create scripts used for debugging that you'll disable once you run scraping sessions in a production environment.
So far we have explained each of the windows in the workbench of screen-scraper. Here we would like to make you aware of a few other windows that you will likely come across in your work with screen-scraper.
The breakpoint window opens when the scraping session runs into a session.breakpoint method call in a script. It is a very effective tool when trouble shooting your scrapes.

The value of any variable can be edited here by double clicking on it, changing it, and deselecting or hitting enter.
The value of any variable can be edited here by double clicking on it, changing it, and deselecting or hitting enter.
This feature is only available to Professional and Enterprise editions of screen-scraper.
At times in developing a scraping session a particular scrapeable file may not be giving you the results you're expecting. Even if you generated it from a proxy session parameters or cookies may be different enough that the response from the server is very different than what you were anticipating, including even errors. Generally in cases like this the best approach is to compare the request produced by the scrapeable file in the running scraping session with the request produced by your browser in the proxy session. That is, ideally your scraping session mimics as closely as possible what your web browser does.
The Compare Last Request and Proxy Transaction window facilitates just such a comparison. I can be accessed in the last request tab of the scrapreable file. After clicking the Compare Last Request and Proxy Transaction button, you will be prompted to select the proxy transaction to which the request should be compared. Simply navigate to the proxy session that it is connected to and select the desired transaction and the window will open.

The screen has four tabs to aid in comparing transaction and request: URL, POST data, Cookies, and Headers. Parameters in any of these areas can be controlled using the scrapeableFile object and its methods.
The DataSet window displays the values matched by the extractor tokens. It can be view in two basic ways:
The DataSet window has two rendering styles. The default is grid view, but you can switch between views using the button at the top of the screen (after view as:).

The names of the columns correspond to the tokens that matched data in the most recent scrapeable file's response. The one addition is the Sequence column that is used by screen-scraper to identify the order in which the matches occurred on the page.
If a column is not showing up for an extractor token it is because that token does not match anything in any of the data records.

This view can be a little easier for viewing the matched data in data record groups.
The regular expressions that you can select for extractor tokens are stored in screen-scraper and can be edited in the Regular Expressions Editor window. The window is accessed by selecting Edit Regular Expressions from the menu.
This can be helpful if you have a regular expression that you use regularly. You can also edit the provided regular expressions though we encourage you not to do so without good reason. These regular expressions have been tested over time and updated when required; they are very stable expressions.

Listed regular expressions can be edited by double clicking in the field that you would like to edit.
One of the most powerful features of screen-scraper is its built-in scripting engine. Through scripting web sites can be crawled in a very dynamic way. Scripting also allows you to insert business logic, clean and normalize data, and write data out to external repositories, such as files and databases. This section of the documentation will familiarize you with scripting in screen-scraper, as well as cover specifics on the various scripting languages that screen-scraper supports.
Before reading through this section you might find it helpful to first read the section on using scripts with the scraping engine. Also, if you haven't done so already, we'd highly recommend going through our first few tutorials, which provides several examples of scripting in screen-scraper.
Because screen-scraper internally uses Java it is important that file paths follow the requirements of Java. That is that file paths follow the Unix/Linux structure (e.g., /usr/local/file.txt). If you are working on a machine that follows these conventions then it will not look any different to you; however, if you are working on a Windows machine this is an important difference to keep in mind.
Windows uses the backslash (\) as a file delimiter but Java uses it as an escape character. That means that on a Windows machine you need to pay closer attention to file paths as they will look a little different.
Windows file paths should use either a forward slash (/) or two backslashes (\\) to delimit file paths.
screen-scraper uses the BeanShell library to allow for scripting in Java. If you've done some programming in C or JavaScript you'll probably find BeanShell's syntax familiar. Documentation for BeanShell is excellent, and we'd recommend referring to it as you program.
Interpreted Java is just a phrase used to mean that it is java that does not require being compiled.
See the using scripts and API pages for details on objects and methods that you can make use of in a script. We also use Interpreted Java in all of our tutorials, which should get you familiar with how it's used in screen-scraper.
It is possible to access Java libraries in screen-scraper. See adding Java libraries for more details.
We use Java in the screen-scraper tutorials but if you would like to learn more about Java you can look for tutorials online. The following are some good Java resources:
This scripting language is not available by default any more. To use it you will need to edit the AllowUnstableWindowsFeatures in the screen-scraper.properties file.
Writing scripts in JScript gives you the familiarity of a widely used language, while still providing access to commonly useed Windows libraries. Using JScript within screen-scraper can only be done on a Windows platform, and requires that the JScript runtime be installed. The chances are good that you've already got the JScript runtime on your system.
screen-scraper will automatically detect if the JScript runtime is installed, which you can see by selecting a script from the objects tree in the workbench and clicking on the Language drop-down list. If you don't see JScript in the list then the runtime needs to be installed.
If you do not have JScript runtime on your system you can download it from Microsoft's script downloads page.
Please be aware that because of a bug in the third-party library that allows screen-scraper to integrate with the Microsoft Scripting Engine problems can occur if multiple JScript scripts are run simultaneously. If you're using the professional edition of screen-scraper and plan on running multiple scraping sessions simultaneously you should use Interpreted Java, JavaScript, or Python as a scripting language.
Because screen-scraper uses the native JScript engine, all Active X objects installed on the computer (such as ADO or the FileSystemObject) can be accessed. Additionally, all of the objects mentioned on the Using scripts and API pages are also available.
Java classes can also be instantiated within a script using the CreateBean function. For example, the following script will instantiate a RunnableScrapingSession and run it:
Mozilla's Rhino scripting engine is used by screen-scraper to allow scripts to be written in JavaScript. Documentation for Rhino is sparse, but the interpreter does adhere strictly to the established ECMAScript standard, so just about any reference on JavaScript could be referred to. If you try writing scripts using JavaScript, and run into difficulties (because of lack of documentation), you may want to consider using Interpreted Java instead, which has very similar syntax and provides significantly better documentation. If you've worked with client-side JavaScript in web programming, you'll probably be comfortable using JavaScript in screen-scraper.
These must be prefaced with the Packages keyword.
This scripting language is not available by default any more. To use it you will need to edit the AllowUnstableWindowsFeatures in the screen-scraper.properties file.
screen-scraper uses ActiveState's ActivePerl library for scripts written in Perl. Using Perl within screen-scraper can only be done on a Windows platform, and requires that the ActivePerl runtime be installed.
screen-scraper will automatically detect if the ActivePerl runtime is installed, which you can see by selecting a script from objects tree in the workbench and clicking on the Language drop-down. If you don't see Perl in the list then the runtime needs to be installed.
The ActivePerl runtime can be downloaded from ActiveState's download page for free.
Java classes can be instantiated within a script using the CreateBean function. For example, the following script will instantiate a RunnableScrapingSession for the "Weather" scraping session (which is found in the default screen-scraper installation) and run it:
The Jython interpreter is used by screen-scraper to for scripting in Python. Jython is a very fast interpreter, and we'd recommend using it if you're familiar with the Python programming language.
Importing your externally-compiled classes is as easy as placing them in the ./lib/ext folder of your installation. The Jython interpreter will automatically include that folder on your PythonPath.
The generator objects are implemented in Jython and the folders lib/ext, lib/jython-lib, and lib/jython-lib/site-packages are included in python's system path.
When scripting in Python all of the standard Java classes can be used. Classes must be imported using the Java package hierarchy of screen-scraper, which is also required if you'd like to create one of screen-scraper's RunnableScrapingSession objects. Here's an example that will run a scraping session called "Weather":
Notice that before the RunnableScrapingSession class can be used it first must be imported.
This scripting language is not available by default any more. To use it you will need to edit the AllowUnstableWindowsFeatures in the screen-scraper.properties file.
If you've programmed in Visual Basic or Active Server Pages you should find scripting in screen-scraper to be similar. Using VBScript within screen-scraper can only be done on a Windows platform, and requires that the VBScript runtime be installed. The chances are good that you've already got the VBScript runtime on your system.
screen-scraper will automatically detect if the VBScript runtime is installed, which you can see by selecting a script from the objects tree in the workbench and clicking on the Language drop-down list. If you don't see VBScript in the list then the runtime needs to be installed.
If you do not have VBScript runtime on your system you can download it from Microsoft's script downloads page.
Please be aware that because of a bug in the third-party library that allows screen-scraper to integrate with the Microsoft Scripting Engine problems can occur if multiple VBScript scripts are run simultaneously. If you're using the professional edition of screen-scraper and plan on running multiple scraping sessions simultaneously you should use Interpreted Java, JavaScript, or Python as a scripting language.
Because screen-scraper uses the native VBScript engine, all Active X objects installed on the computer (such as ADO or the FileSystemObject) can be accessed. Additionally, all of the objects mentioned on the using scripts and API pages are also available.
Java classes can also be instantiated within a script using the CreateBean function. For example, the following script will instantiate a RunnableScrapingSession and run it:
The web interface is only available for enterprise edition users of screen-scraper.
The screen-scraper web interface allows you to administer aspects of the scraping process. This includes monitoring running scraping sessions, importing and exporting scraping sessions, and scheduling scraping sessions to be run on a periodic basis.

When screen-scraper is running in server mode, you can access the web interface on your local machine at the following URL: http://localhost:8779/.
If you've changed the Web/SOAP Server port in the workbench or the SOAPPort in the screen-scraper.properties file, you'll need to use the port you designated.
Depending on the operating system you're running, instead of localhost, you may need to use 127.0.0.1 or the IP address of the machine.
If a scraping session or script with the same name already exists in this instance of screen-scraper then script overwriting properties will determine which one is discarded. If a script cannot be overwritten then a warning message will inform you of that the script that was trying to import was discarded.
If screen-scraper is running when you copy the files into the import folder they will be imported and hot-swapped in the next time a scraping session is invoked. They will also be imported if you start or stop screen-scraper.
The web interface settings can be opened by clicking on the settings button in the upper-right corner of the screen.

If this value is blank, 0, or negative, the scraping session will not time out.
Flagged scrapes are highlighted in red in the run/running tab.
This tab displays all scraping sessions loaded into the current instance of screen-scraper. It will display basic information on scraping sessions that are currently running, as well as scraping sessions that have run in the past. It also allows you to start and schedule scraping sessions.
The runnable tab will display all of the scraping sessions listed alphabetically by name, and the messages from the most recently started instance of the scrape.

If the scraping session is not currently running it shows is how long is took to run last time it was run.
If the scraping session is not currently running it show the amount of time it took to run two times ago.
This tab displays information on scraping sessions that are either currently running or have run in the past. You can use this table to compare run times, the number of records scraped, and also to monitor scraping session logs. If scraping sessions have timed out (see settings) the stop button will gray and the status will change to interupted. If a script has flagged a fatal error (see setFatalErrorOccurred) then the error cell will display in red for that scrape.
Scrapes can be ordered in ascending and descending order using any of the fields. This is done by clicking on the column header that you want to sort by.

Removing records for scraping sessions that have run doesn't remove the scraping sessions themselves, just the records related to the time when they were run.
Removing records for scraping sessions that have run doesn't remove the scraping sessions themselves, just the records related to the time when they were run.
If the scraping session is not currently running it shows is how long is took to run last time it was run.
If the scraping session is not currently running it show the amount of time it took to run two times ago.
On this tab you can manage scraping sessions that have been scheduled to be run. The columns can be sorted by clicking on the column headers.

If this value is 0 or a negative number, the scraping session will not time out.
If the run of the scraping session is disabled, it will not run even if it's scheduled to do so.
It can be very helpful to have scraping sessions run automatically or on an on going basis. The web interface makes this simple allowing you to schedule and manage multiple scrapes in a single location.

If this value is blank, 0, or negative, the scraping session will not time out.

If these boxes are left blank, the scraping session will run once and not be re-scheduled.

Flagged scrapes are highlighted in red in the run/running tab.
When writing scripts within screen-scraper, there are a number of objects and methods available to you. The Using Scripts page provides an overview of working with scripts, where this page provides details on specific objects and methods you'll use when scripting within screen-scraper.
The API documentation emphasizes Interpreted Java as Java is the language in which screen-scraper proper is written. That should not deter you from using whatever language you desire; all the methods are available in what ever language you choose.
The examples given here assume you're using Interpreted Java as the scripting language, but there should be very little difference in syntax if you decide to use another language. For example, if you're scripting in VBScript, you would simply omit the semi-colon at the end of each line, and for methods that don't return a value you would precede them with the VBScript keyword Call (either that, or omit the parentheses around the method parameters).
The screen-scraper, internal API has been divided into three groups for convenience.
The two main groups are the scraping engine and the proxy server. The various objects available in these sections are exclusive to running screen-scraper for in one of these two ways. The one exception is the RunnableScrapingSession which has been grouped with the scraping engine simply because it is unlikely to be needed or used with the proxy server.
The utilities are available to scripts run in either the scraping engine or the proxy server and have since been separated from both. These represent classes that we have written to simplify some common tasks that are performed with retrieved data.
There are many additional classes that are available through Java Libraries that we did not create/modify that are especially worthy of note. Regardless of the language that you are using to program in screen-scraper you can have access to these.
There are a few other APIs to be aware of. They are particular to dealing with screen-scraper in certain ways or certain versions. Make sure that you understand the implications of using these APIs before you start playing with them.
The scraping engine is the backbone of screen-scraper and provides four built-in objects. These objects are: session, scrapeableFile, dataSet, and dataRecord. We have also included the RunnableScrapingSession class as it best pertains to the engine.
For details on which objects are available to scripts in the context of a scrape see the variable scope section of the documentation.
The dataRecord object is populated using the names of tokens from extractor patterns.
This object gives access to the most recently extracted data record. This will most likely only be used in scripts that get accessed after each time an extractor pattern is applied. This object simply extends Hashtable (documentation on its methods can be found in Java's documentation).
The dataRecord is populated using the token names in the extractor patterns. You'll find a few of the most commonly used methods below. DataRecord objects can also be created from scratch, and subsequently added to DataSet objects using the addDataRecord method.
See example usage: Iterate over DataSets & DataRecords.
Create a new DataRecord object.
This method does not receive any parameters.
Returns DataRecord object.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
com.screenscraper.common.DataRecord
See additional example usage: Iterate over DataSets & DataRecords.
Get the value of a DataRecord field.
Returns the value associated with the specified key. Usually it will be a string but, if you have manually added fields, it can be an integer, boolean, long, or other object.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Add a new field to the DataRecord or update the value of an existing field.
Returns the value previously associated with the specified key. If the key did not exist then it will return null.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
See additional example usage: Iterate over DataSets & DataRecords.
Remove a field from the DataRecord.
Returns the value previously associated with the specified key. If the key did not exist then it will return null.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
The dataSet object holds all data records extracted by an extractor pattern after it has been applied as many times as possible to the HTML retrieved by a scrapeable file. A data set is analogous to a result or record set that would be returned from a database query. A data set contains any number of data records, which are analogous to rows in a database.
The dataSet object provides methods to aid in getting at the information that has been gathered.
See example usage: Iterate over DataSets & DataRecords.
Manually create a DataSet.
Returns DataSet object.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
com.screenscraper.common.DataSet
See additional example usage: Iterate over DataSets & DataRecords.
Add a DataRecord to a DataSet.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
See additional example usage: Iterate over DataSets & DataRecords.
Remove all DataRecord objects from the DataSet.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
See additional example usage: Iterate over DataSets & DataRecords.
Remove a DataRecord from the DataSet.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve a field's value in a data set based on another field.
Returns the value in the returned column, usually a string (unless records have been manually added). If no match is found, null is returned.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Get a single piece of data held by a DataRecord in the DataSet.
Returns the value associated with the DataRecord identifier. It will be a string unless you have added values to the DataRecord whose values are not strings.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Get all DataRecords in the DataSet.
This method does not receive any parameters.
Returns an ArrayList of DataRecord objects.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
This method is provided as a convenience, the recommended way to iterate over data records in a data set is to use getNumDataRecords and getDataRecord.
Get the character set being applied the scraped data.
This method does not receive any parameters.
Returns the character set applied to the scraped data, as a string. If a character set has not been specified then it will default to the character set specified in settings dialog box.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Get one DataRecord in the DataSet.
Returns a DataRecord (Hashtable object). If there is not a DataRecord at the specified index an error will be thrown.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Get the first non-null value, in a data set, for a given token.
Returns the first non-null value in the column, usually a string (unless records have been manually added). If none is found, null is returned.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Get the number of DataRecords in the DataSet.
This method does not receive any parameters.
Returns the number of DataRecords in the DataSet, as an integer.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Merge data records from two data sets.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Set the character set to be used for rendering dataSet values.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
This will only change the character set on the current data set. If you want it to be changed for all data sets, you would need to change it in the settings dialog box or screen-scraper.properties file.
Get the number of DataRecords in the DataSet.
This method does not receive any parameters.
Returns the number of DataRecords in the DataSet, as an integer.
| Version | Description |
|---|---|
| 6.0.3a | Available for all editions. |
Write DataSet string and integer contents to a file. The fields will be tab-delimited and records hard-return delimited.
Returns void. If the file cannot be written to then an error will be thrown.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This object contains various methods used to log information about a running scraping session to log files, the workbench "Log" tab, and the web interface.
Creates an automatic progress bar and adds it to the progress bars. These progress bars match their progress to a value from a session variable and a list of values. When web messages are output with the webDebug, webInfo, webWarn, or webError methods, a progress bar will be drawn to give a visual representation of the current progress of the scrape.
Note that when using auto progress bars, it is advised to not use any manually monitored ones, as it can cause conflicts. Anytime an auto progress bar has no session variable set for its monitored key, it deletes itself and all children progress bars (including manual ones). As long as you keep that in mind, it should be safe to use both types together.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.31a | Available in enterprise edition. |
| 5.5.43a | Moved from session to log class. |
Watches for all session variables whose keys end with the postfix specified, and will output their values when monitored variables are logged.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.42a | Moved from session to log class. |
Watches for all session variables whose keys begin with the prefix specified, and will output their values when monitored variables are logged.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.42a | Moved from session to log class. |
Adds a specific name and value to be logged with the web messages methods or logMonitoredValues method
The previous value associated with the name, or null if there wasn't one
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.42a | Moved from session to log class. |
Watches the value of a session variable, and will output it each time monitored variables are output
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.42a | Moved from session to log class. |
Adds a new progress bar. If no progress bar exists, this will be set as the root, otherwise it will be the child of the lowest progress bar. When web messages are output with the webDebug, webInfo, webWarn, or webError methods, a progress bar will be drawn to give a visual representation of the current progress of the scrape. The addProgressBarIfNotStopped versions remove the progress bar if the scrape has not been stopped, which is useful for determining when a scrape was stopped.
This method returns a reference to the new progress bar, which can be used to update the current progress
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.31a | Available in enterprise edition. |
| 5.5.43a | Moved from session to log class. |
Appends a status message to be displayed in the web interface.
None
| Version | Description |
|---|---|
| 5.5.32a | Available in Enterprise edition. |
| 5.5.43a | Moved from session to log class. |
Adds a file to the cache. This can be used to add anything to the cache, from a text file to an image that was downloaded, or any other file that would be useful.
A File that represents the cached file.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
Caches the HTML and headers of the scrapeable file. This will include both the request and response headers.
A File that represents the cached file.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
Adds text to the cache. This will create a new text file in the cache and store the given content in it.
A File that represents the cached file.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
Write message to the log.
Returns void.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for all editions. |
When the workbench is running, this will be found under the log tab for the scraping
session. When screen-scraper is running in server mode, the message will get sent to the corresponding .log file found in screen-scraper's log folder. When screen-scraper is invoked from the command
line, the message will get sent to standard out.
Enables caching for this scrape. When caching is enabled, each time a scrapeable file is downloaded it will be saved to the file system. Once the session is completed the cache will be either zipped or the directory renamed, depending on the conditions that were specified when the cache was enabled. Optionally this will save the log files to the cached location, and will save everything from the error.log file that was added while the cache was enabled.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.32a | Renamed from enableCache to enableCaching |
| 5.5.43a | Moved from session to log class. |
Ends the caching for the scrape. This method will be called once all the scripts and files are run/scraped. It can be called in a script to end the caching early (thereby only caching a portion of the scrape). This only deals with saving downloaded content to the file system, not with reading it back in during a scrape.
This method takes no parameters
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.32a | Renamed from endCache to endCaching. |
| 5.5.43a | Moved from session to log class. |
Write message to the log.
Returns void.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for all editions. |
When the workbench is running, this will be found under the log tab for the scraping
session. When screen-scraper is running in server mode, the message will get sent to the corresponding .log file found in screen-scraper's log folder. When screen-scraper is invoked from the command
line, the message will get sent to standard out.
Returns whether or not the cache is enabled for the scrape. When enabled, it simply means that each ScrapeableFile will save the content it downloads from the server to the file system so it can be viewed later, generally for debugging purposes.
This method takes no parameters
Returns true if caching is currently enabled for this session
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.32a | Available enterprise and professional editions (Returns false in basic edition, but doesn't throw an Exception). Renamed from getCacheEnabled to getCachingEnabled. |
| 5.5.43a | Moved from session to log class. |
Returns the progress bar specified. If the index if given, returns the progress bar at that index (0 is the root, 1 is the first child, etc...). If the title is given, returns the most recently added progress bar with the given title
The ProgressBar indicated, or null if none was found matching the required criteria
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.31a | Available in enterprise edition. |
| 5.5.43a | Moved from session to log class. |
Write message to the log.
Returns void.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for all editions. |
When the workbench is running, this will be found under the log tab for the scraping
session. When screen-scraper is running in server mode, the message will get sent to the corresponding .log file found in screen-scraper's log folder. When screen-scraper is invoked from the command
line, the message will get sent to standard out.
Write message to the log.
Returns void.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for all editions. |
When the workbench is running, this will be found under the log tab for the scraping
session. When screen-scraper is running in server mode, the message will get sent to the corresponding .log file found in screen-scraper's log folder. When screen-scraper is invoked from the command
line, the message will get sent to standard out.
Logs all the values in a Data Record to the log, with one line per value. If a value in the record is a List, Set, Map, Data Set, Scrapeable File, or Exception, it will have detailed output as well.
This method returns nothing
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
The output from the above call might look something like this:
DataRecord --- A_FLOAT : 3.14159 --- A_LIST : List ------ Element 0 : Value 1 ------ Element 1 : Value 2 ------ Element 2 : Value 3 ------ Element 3 : Set --------- Element : A value --------- Element : More value --------- Element : Other stuff --- A_MAP : Map ------ KEY_1 : 1 ------ KEY_2 : 2 ------ KEY_3 : 3 --- A_SET : Set Logged above as "------ Element 3 : " --- A_STRING : Screen-Scraper --- AN_INT : 5
Logs an Exception, with a full stack trace, at the Error level
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Logs the values of all the currently monitored variables, the progress of the scrape, if known, and puts the message at the top. Also logs any additional values being watched. Logs values at the specified level.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Logs closing values to indicate the scrape is complete and what values were when everything finished. It will log at whatever the highest level logged to was. For instance, if a webWarn had been logged during the scrape, this will log at the warning level.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Logs the Object in a semi intelligent way. For example, Maps are logged as key-value pairs, lists are logged with one element per line, all elements of a set are logged, etc... Some objects will just log their value using String.valueOf() if it isn't a standard type of data set/list
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Logs useful information about the current instance of Screen-Scraper, as well as the Java VM and the General Utility version being used. Information will be logged as an info message in the web interface (when running in server mode) and the log.
This method takes no parameters
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Stops watching for a postfix in session variables
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Stops watching for a prefix in session variables
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Removes a specific name from the manually set values to be logged. Doesn't affect the value of session variables
The previous value associated with the name, or null if there wasn't one
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Stops watching the specified variable
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Moved from session to log class. |
Removes the specified progress bar. The removeProgressBarIfNotStopped version removes the progress bar if the scrape has not been stopped, which is useful for determining when a scrape was stopped.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.31a | Available in enterprise edition. |
| 5.5.43a | Moved from session to log class. |
Write message to the log.
Returns void.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for all editions. |
When the workbench is running, this will be found under the log tab for the scraping
session. When screen-scraper is running in server mode, the message will get sent to the corresponding .log file found in screen-scraper's log folder. When screen-scraper is invoked from the command
line, the message will get sent to standard out.
Logs closing values to indicate the scrape is complete and what values were when everything finished. It will log at whatever the highest level logged to was. For instance, if a webWarn had been logged during the scrape, this will log at the warning level. When running in Professional edition, this simply outputs to the log.
Using this method is preferred over logMonitoredValuesClose (which only logs to the log), because if at a later point the scrape is run in server mode for enterprise edition, a useful message is output in the web interface without needing to modify the scrape.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
Logs a debug message to the web interface status message area. Uses the message header as the top of the message, and then logs all currently monitored session variables underneath as well as the current progress (if known) of the scrape. Also outputs the message to the log. When running in Professional edition, this simply outputs to the log.
Using this method is preferred over logMonitoredValues (which only logs to the log), because if at a later point the scrape is run in server mode for enterprise edition, a useful message is output in the web interface without needing to modify the scrape.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
Logs an error message to the web interface status message area. Uses the message header as the top of the message, and then logs all currently monitored session variables underneath as well as the current progress (if known) of the scrape. Also outputs the message to the log. When running in Professional edition, this simply outputs to the log.
Using this method is preferred over logMonitoredValues (which only logs to the log), because if at a later point the scrape is run in server mode for enterprise edition, a useful message is output in the web interface without needing to modify the scrape.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
Logs an info message to the web interface status message area. Uses the message header as the top of the message, and then logs all currently monitored session variables underneath as well as the current progress (if known) of the scrape. Also outputs the message to the log. When running in Professional edition, this simply outputs to the log.
Using this method is preferred over logMonitoredValues (which only logs to the log), because if at a later point the scrape is run in server mode for enterprise edition, a useful message is output in the web interface without needing to modify the scrape.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
Logs a warning message to the web interface status message area. Uses the message header as the top of the message, and then logs all currently monitored session variables underneath as well as the current progress (if known) of the scrape. Also outputs the message to the log. When running in Professional edition, this simply outputs to the log.
Using this method is preferred over logMonitoredValues (which only logs to the log), because if at a later point the scrape is run in server mode for enterprise edition, a useful message is output in the web interface without needing to modify the scrape.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
| 5.5.43a | Moved from session to log class. |
This is a class that can be instantiated within a script in order to run a scraping session.
Also see:
The Maximum number of concurrent running scraping sessions in the settings dialog box will control how many scraping sessions can be run simultaneously.
Initiates a RunnableScrapingSession object using the name of an existing scraping session.
Returns a RunnableScrapingSession. On failure an error will be thrown.
| Version | Description |
|---|---|
| 5.0 | inheritHttpState added as optional parameter. |
| 4.5 | Available for professional and enterprise editions. |
com.screenscraper.scraper
Retrieve the name of the scraping session in the runnableScrapingSession.
This method does not receive any parameters.
Returns a string with the name of the scraping session.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Get the timeout of the session in the runnableScrapingSession.
This method does not receive any parameters.
Returns a integer representing the timeout length in minutes.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Retrieve the the value of a session variable. This method should be called after scrape method has returned.
Returns the value of the session variable: object, boolean, int, string, etc. If the variable doesn't exists it returns null.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Run the session scraping.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
The default is for the script to continue executing without waiting for the scraping session to finish. You can use setDoLazyScrape to force the script to wait until the scape finishes before continuing the script.
Indicate whether or not the scraping session should run concurrently with (at the same time as) other scraping sessions. The default for doLazyScrape is true.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
We recommend not setting this value to false! When running scraping sessions in the workbench, it will cause the interface to freeze up until sessions have completed.
If you'd like to run multiple scraping sessions serially (one after another), the best option is to set the Maximum number of concurrent running scraping sessions to 1 in the settings window.
Sets the timeout of the session. That is, after the given number of minutes have passed the session will automatically terminate.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method must be called before scrape.
Set the value of a session variable.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
The scrapeableFile object refers to the current file being requested from a given server. It houses both the request for a file and response and can be manipulated to meet any necessary requirements: GET and POST parameters, referer information, cookies, FILE parameters, HTTP headers, characterset, and such.
Dynamically adds a GET parameter to the URL of the current scrapeable file. If a parameter with the given sequence already exists, it will be replaced by the one created from this method call. Calling this method is the equivalent in the workbench of adding a parameter under the "Parameters" tab, and designating the type as GET. Once the scraping session is completed the original HTTP parameters (those under the "Parameters" tab in the workbench) will be restored.
None
| Version | Description |
|---|---|
| 5.5.32a | Available in Professional and Enterprise editions. |
Add an HTTP header to be sent along with the request.
Returns void. If you are not using enterprise edition it will throw an error.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise edition. |
| 4.5 | Available for enterprise edition. |
In certain rare cases it may be necessary to explicitly add a custom header of the POST data of an HTTP request. This may be required in cases where a site is using AJAX, and the POST payload of a request is sent as XML (e.g., using the setRequestEntity method). This method must be invoked before the HTTP request is made (e.g., "Before file is scraped" for a scrapeable file).
Dynamically add an HTTPParameter to the current scrapeable file.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
The HTTPParameter constructor is as follows: HTTPParameter( String key, String value, int sequence, String type ). Valid types for the constructor are GET, POST, and FILE. Calling this method will have no effect unless it's invoked before the file is scraped.
Dynamically adds a POST parameter to the existing set of POST parameters. If a parameter with the given sequence already exists, it will be replaced by the one created from this method call. If the method call is used that doesn't take a sequence, the new POST parameter will carry a sequence just higher than the highest existing sequence. Calling this method is the equivalent in the workbench of adding a parameter under the "Parameters" tab, and designating the type as POST. Once the scraping session is completed the original HTTP parameters (those under the "Parameters" tab in the workbench) will be restored.
None
| Version | Description |
|---|---|
| 5.5.32a | Available in Professional and Enterprise editions. |
Manually apply an extractor pattern to a string.
Returns DataSet on success. Failures will be written out to the log as errors.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
An example of how to manually extract data is available.
Manually retrieve the value of a single extractor token.
Returns the match from the last data record, as a string, on success. On failure it returns null and writes a error to the log.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
If you want it to be from the first data record you could use getDataRecord.
Gets the ASPX .NET values from the string. The standard values are __VIEWSTATE, __EVENTTARGET, __EVENTVALIDATION, and __EVENTARGUMENT. Values will be stored in the returned DataRecord as ASPX_VIEWSTATE, ASPX_EVENTTARGET, etc...
A DataRecord object with each ASPX name as ASPX_[NAME] mapped to it's value. Note that when onlyStandard is false, any parameter that starts with the name __ will be returned in this DataRecord
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Retrieve the authentication expectation of the request.
This method does not receive any parameters.
Returns whether the scrapeable file expects to have to authenticate and so will send the information initially instead of waiting for the request for it, as a boolean.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Get the character set being used in the page response rendering.
This method does not receive any parameters.
Returns the character set applied to the scraped page, as a string. If a character set has not been specified then it will default to the character set specified in settings dialog box.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
If you are having trouble with characters displaying incorrectly, we encourage you to read about how to go about finding a solution using one of our FAQs.
Retrieve contents of the response.
This method does not receive any parameters.
Returns contents of the last response, as a string. If the file has not been scraped it will return an empty string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve the POST payload type being used to interpret the page. This can be important with scraping some site's implementation of AJAX, where the payload in explicitly set as xml.
This method does not receive any parameters.
Returns the content type, as a string (e.g., text/html or text/xml).
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Retrieve the POST data.
This method does not receive any parameters.
Returns the POST data for the scrapeable file, as a string. If called after the file has been scraped the session variable token will be resolved to their values; otherwise, the tokens will simply be removed from the string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Get the URL of the file.
This method does not receive any parameters.
Returns the URL of the scrapeable file, as a string. If called after the file has been scraped the session variable tokens will be resolved to their values; otherwise, the tokens will simply be removed from the string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Indicates whether or not the most recent extractor pattern application timed out.
None
| Version | Description |
|---|---|
| 5.5.36a | Available in all editions. |
Determine whether or not the contents of this response are being forced to be recognized as non-binary.
This method does not receive any parameters.
Returns true if the scrapeable file is being forced to be treated as non-binary; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Gets the value of the header in the response of the scrapeable file, or returns null if it couldn't be found
The value of the header, or null if not found
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Gets the header section of the HTTP Response
This method takes no parameters
A String containing the HTTP Response Headers
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Gets the headers of the HTTP Response as a map, and returns them.
This method takes no parameters
A Map from header name to header value for the response headers.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Indicates whether or not the most recent attempt to tidy the HTML failed.
None
| Version | Description |
|---|---|
| 5.5.36a | Available in all editions. |
Indicates whether or not the maximum attempts to request a given scrapeable file were reached.
None
| Version | Description |
|---|---|
| 5.5.36a | Available in all editions. |
Retrieve the kilobyte limit for information retrieved by the scrapeable file, any additional information will not be retrieved.
This method does not receive any parameters.
Returns the current kilobyte limit on the response, as an integer.
| Version | Description |
|---|---|
| 5.0 | Add for professional and enterprise editions. |
Get the name of the scrapeable file.
This method does not receive any parameters.
Returns the name of the scrapeable file, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve the non-tidied HTML of the scrapeable file.
This method does not receive any parameters.
Returns the non-tidied contents of the scrapeable file, as a string. On failure it returns null.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
By default non-tidied html is not retained. For this method to return anything other than null you must use setRetainNonTidiedHTML to force non-tidied html to be retained.
Gets an array of strings containing the redirect URL's for the current scrapeable file request attempt.
This method does not receive any parameters.
Returns the array of strings; may be empty.
| Version | Description |
|---|---|
| 6.0.24a | Available in Professional and Enterprise editions. |
Determine if the scrapeable file is set to retain non-tidied html.
This method does not receive any parameters.
Returns boolean flag for non-tidied contents being retained.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Returns the retry policy. Note that in any 'After file is scraped' scripts this is null
This method takes no parameters.
The Retry Policy that will be used by this scrapeable file
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Determine the HTTP status code sent by the server.
This method does not receive any parameters.
Returns integer corresponding to the HTTP status code of the response.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Retrieve the name of the user agent making the request.
This method does not receive any parameters.
Returns the user agent, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Determine if an input or output error occurred when requesting file.
This method does not receive any parameters.
Returns true if an error has occurred; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
This method should be run after the scrapeable file has been scraped.
Determine whether any extractor patterns associated with the scrapeable file found a match.
This method does not receive any parameters.
Returns boolean corresponding to whether any extractor pattern matched in the scrapeable file.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Remove all of the HTTP parameters from the current scrapeable file.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Remove an HTTP header from a scrapeable file.
Returns void.
| Version | Description |
|---|---|
| 5.0.5a | Introduced for enterprise edition. |
Dynamically removes an HTTPParameter. The order of the remaining parameters are adjusted immediately.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
| 5.5.32a: Added method call that takes a String. | Available for Professional and Enterprise editions. |
If calling this method more than once in the same script, when used in conjunction with the addHTTPParameter method, it is important to keep track of how the list is reordered before calling either method again.
Calling this method will have no effect unless it's invoked before the file is scraped.
This method can be used for both GET and POST parameters.
Resequences an HTTP parameter.
None
| Version | Description |
|---|---|
| 5.5.32a | Available in Professional and Enterprise editions. |
Resolves a relative URL to an absolute URL based on the current URL of this scrapeable file.
Returns string containing the complete url to the file. On failure it will return the relative path and an error will be written to the log.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Write non-tidied contents of the scrapeable file response to a text file.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method must be called before the file is scraped.
Because the response header are also saved in the file, if the file is anything except a text file it will not be valid (e.g. images, pdfs).
Save the file returned from a scrapeable file request.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
This method must be called from a scrapeable file before the file is scraped. Do not call this method from a script which is invoked by other means such as after an extractor pattern match or from within another script.
It is preferable to use downloadFile; however, at times you may have to send POST parameters in order to access a file. If that is the case, you would use this method.
This method cannot save local file requests to another location.
Set the authentication expectation of the request.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Set the character set used in a specific scrapeable file's response renderings. This can be particularly helpful when the page renders characters incorrectly.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
This method must be called before the file is scraped.
If you are having trouble with characters displaying incorrectly, we encourage you to read about how to go about finding a solution using one of our FAQs.
Set POST payload type. This is particularly helpful with scraping some site's implementation of AJAX, where the payload in explicitly set as xml.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method must be called before the file is scraped.
This method is usually used in connection with setRequestEntity as that method specifies the content of the POST data.
Set content type header to multipart/form-data.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method must be called before the file is scraped.
Occasionally a site will expect a multi-part request when a file is not being sent in the request.
If you include a file upload parameter under the parameters tab of the scrapeable file the request will automatically be multi-part.
Set whether or not the contents of this response should be forced to be treated as non-binary. Default forceNonBinary value is false.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
This is provided in the case where screen-scraper misidentifies a non-binary file as a binary file. It doesn't happen often but is possible.
Determines whether or not a POST request should be forced.
Returns void.
| Version | Description |
|---|---|
| 6.0.14a | Available in Professional and Enterprise editions. |
Sets the request type to use.
ScrapeableFile.RequestType is an enum with the following options as values
If the method sets the request to one of those types, all paramenters set as GET in the paramenters tab will be appended to the url (like normal) and all parameters set as POST parameters will be used to buld the request entity. If there are POST values on a type that doesn't support a request entity an exception will be thrown when the request is issued.
Returns void.
| Version | Description |
|---|---|
| 6.0.55a | Available in Professional and Enterprise editions. |
Overwrite the content of the "last response"
Returns void.
This method must be called from an extractor pattern before the pattern is run.
Limit the amount of information retrieved by the scrapeable file. This method can be useful in cases of very large responses where the desired information is found in the first portion of the response. It can also help to make the scraping process more efficient by only downloading the needed information.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Add for professional and enterprise editions. |
This method must be called before the file is scraped.
Set referer HTTP header.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method must be called before the file is scraped.
Set POST payload data. This is particularly helpful with scraping some site's implementation of AJAX, where the payload in explicitly set as xml.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method must be called before the file is scraped.
This method is usually used in connection with setContentType as that method specifies the content of the POST data.
Though you can set plain text POST data using this method it is preferable to use the addHTTPParameter method for this task.
Set whether or not non-tidied HTML is to be retained for the current scrapeable file.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
If, after the file is scraped, you want to be able to use getNonTidiedHTML this method has to be called before the file is scraped.
Sets a Retry Policy that will be run to check if a page should be re-downloaded or not. The policy will be checked after all the extractors have run, and will check for an error on the page based on a set of conditions. If the policy shows an error on the page, it can run scripts or other code to attempt to remedy the situation, and then it will rescrape the file.
The file will be re-downloaded without rerunning any of the scripts that run before the file is downloaded, and before any of the scripts marked to run after the file is scraped. If there is any change that needs to be made to session variables/headers, etc... they should be made in the script or runnable that will be executed. Also, the policy can specify that session variables should be restored to their previous values before the file is rescraped. If it does, they will be reset after the error checking portion of the policy but before the policy runs the code to make changes before a retry.
The retry policy should be set in a script run 'Before file is scraped', but can also be set by a script on an extractor pattern. It it is set on an extractor pattern, session variables will not be restored if the retry is required
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Explicitly state the user agent making the request.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method must be called before the file is scraped.
Determine if an error occurred with the request. Errors are considered to be server timeouts as well as any status code outside of the range 200-399.
This method does not receive any parameters.
Returns true for server timeouts as well as any status code outside of the range 200-399; otherwise, it returns false.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
This method must be called after the file is scraped.
If you want to know what the status code was you can use getStatusCode.
This object refers to the current scraping session that is running. To make the methods a little easier to sort through they have been grouped into related methods. The groups have been named to ease in finding them when they are needed.
The following methods are provided to aid you in setting up an anonymous scraping session. If you are using your own server proxy pool you will use the methods to allow screen-scraper to interact with and manage your proxy pool. If you are using automatic anonymization then the only method you will use is currentProxyServerIsBad as screen-scraper will manage the servers using the anonymization settings from your setup.
See an example of Anonymization via Manual Proxy Pools.
Remove proxy server from proxy pool. This is only used with anonymization and indicates that one server in the pool is bad and should be removed.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
If you are using automatic anonymization or manual proxy pools, a new proxy server will be created as a result of the method call.
When checking if a request you have made is invalid it is best not to rely on the HTTP status code (eg. 404) alone as the status codes are not always accurate. It is recommended that you also scrape a known string (eg. "Not found") from the response HTML that validates the status code.
Get the current proxy server from the proxy server pool.
This method does not receive any parameters.
Returns the current proxy server being used.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Holds the proxy server pool object that allows proxies to be cycled through.
Returns true if there is an available proxy server pool.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Determine whether proxies are set to be terminated when the scrape ends.
This method does not receive any parameters.
Returns true if a proxy will be terminated; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Determine whether proxies are being used from proxy pool.
This method does not receive any parameters.
Returns true if a proxy pool is being used; otherwise, it returns false.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Associate a proxy pool with a scraping session.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Manually set the outcome of proxies when the scrape ends.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Determine if proxies from a proxyServerPool be used when making scrapeable file request.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
If you are already going through a proxy server, screen-scraper must be told the credentials in order to get out to the internet. These methods are all provided to manually tell screen-scraper how to get through your external proxy.
If you always go through the same external proxy you would probably want to set the credentials in screen-scraper's proxy settings so that you don't have to specify them in all of your scrapes.
Retrieve the external NT proxy domain.
This method does not receive any parameters.
Returns the external NT domain, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Retrieve the external NT proxy host.
This method does not receive any parameters.
Returns the external NT host, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Retrieve the external NT proxy password.
This method does not receive any parameters.
Returns the external NT password, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Retrieve the external NT proxy username.
This method does not receive any parameters.
Returns the external NT username, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Retrieve the external proxy host.
This method does not receive any parameters.
Returns the external host, as a string.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Retrieve the external proxy password.
This method does not receive any parameters.
Returns the external password, as a string.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Retrieve the external proxy port.
This method does not receive any parameters.
Returns the external port, as a string.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Retrieve the external proxy username.
This method does not receive any parameters.
Returns the external username, as a string.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Manually set external NT proxy domain.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external NT proxy settings.
If you are using NTLM (Windows NT) authentication you'll need to designate settings for both the standard external proxy as well as the external NT proxy.
Manually set external NT proxy host/domain.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external NT proxy settings.
If you are using NTLM (Windows NT) authentication you'll need to designate settings for both the standard external proxy as well as the external NT proxy.
Manually set external NT proxy password.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external NT proxy settings.
If you are using NTLM (Windows NT) authentication you'll need to designate settings for both the standard external proxy as well as the external NT proxy.
Manually set external NT proxy username.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external NT proxy settings.
If you are using NTLM (Windows NT) authentication you'll need to designate settings for both the standard external proxy as well as the external NT proxy.
Manually set external proxy host/domain.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external proxy settings.
Manually set external proxy password.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external proxy settings.
Manually set external proxy port.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external proxy settings.
Manually set external proxy username.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
If you are using this method on all of your scripts you might want to set it in screen-scraper's external proxy settings.
Use of log is a great tool to ensure that your scrapes are working correctly as well as troubleshooting problems that arise. Though logging large amounts of information may slow down a scrape, the best way around this is not to remove log writing requests but rather change the verbosity of the logging when running the scrape in a production environment. If you do this, know that you make it harder to troubleshoot some problems should they arise.
The number of methods provided is merely to enhance your ability to log information according to importance.
Get the name of the current log file.
This method does not receive any parameters.
Returns the name of the log file, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method can be very helpful when screen-scraper is running in server mode and you are tracking the log where the scrape of a record is located, or for tracking the location of errors in larger scrapes.
Write message to the log.
Returns void.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for all editions. |
When the workbench is running, this will be found under the log tab for the scraping session. When screen-scraper is running in server mode, the message will get sent to the corresponding .log file found in screen-scraper's log folder. When screen-scraper is invoked from the command line, the message will get sent to standard out.
Write current date and time to log (at most verbose level). It is formatted to be human readable.
This method does not receive any parameters.
Returns void. If an error occurs, an error will be thrown.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Write current time to log (at most verbose level). The time is formatted to be human readable.
This method does not receive any parameters.
Returns void. If an error occurs, an error will be thrown.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Write message to the log, at the the debug level (most verbose).
Returns void.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for professional and enterprise editions. |
Write scrape run time to the log (at most verbose level). It is formatted to be human readable, including breaking it into days, hours, minutes, and seconds.
This method does not receive any parameters.
Returns void. If an error occurs, an error will be thrown.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Write message to the log, at the the error level (least verbose).
Returns void. If an error occurs, an error will be thrown.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for professional and enterprise editions. |
Write message to the log, at the the info level (second most verbose).
Returns void. If an error occurs, an error will be thrown.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for professional and enterprise editions. |
Write all session variables to log.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Write message to the log, at the the warn level (third most verbose).
Returns void. If an error occurs, an error will be thrown.
| Version | Description |
|---|---|
| 5.5 | Now accepts any Object as a message |
| 4.5 | Available for professional and enterprise editions. |
These methods are used in connection with the web interface of screen-scraper. Their use will provide the interface with more detailed information regarding the state of a running scrape. If you are not running the scrapes using the web interface then these methods are not particularly helpful to you.
As the web interface is an enterprise edition feature, these methods are only available in enterprise edition users.
Add to the value of duplicate records scraped. (As opposed to new or error records.)
Returns void.
| Version | Description |
|---|---|
| 7.0 | Available for enterprise edition. |
Add to the value error records. (As opposed to duplicate or new records.)
Returns void.
| Version | Description |
|---|---|
| 7.0 | Available for enterprise edition. |
Add to the value of new records scraped. (As opposed to duplicate or error records.)
Returns void.
| Version | Description |
|---|---|
| 7.0 | Available for enterprise edition. |
Add to the value of number of records scraped.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Append an error message to any existing error messages.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Get the current error message.
This method does not receive any parameters.
Returns current error message, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Determine the fatal error status of the scrape.
This method does not receive any parameters.
Returns whether a fatal error has occurred, as a boolean .
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Get the number of records that have been scraped.
This method does not receive any parameters.
Returns number of records scraped, as a integer.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Reset the count on the number of scraped records.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Set the current error message.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Set the fatal error status of the scrape.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Set the number of records that have been scraped.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Add a runnable that will be executed at the given time.
Note: session.addEventCallback is automatically executed at a priority of 0.
Returns void.
| Version | Description |
|---|---|
| 6.0.55a | Introduced for pro and enterprise editions. |
The EventFireTime is an interface which defines the methods that a fire time must have and so the addEventCallback method can take different types of fire times.
A number of different types of classes based on this interface have been defined for you which call out the various parts of a scrape that you can add event handlers to. Those are defined below.
| Version | Description |
|---|---|
| 6.0.55a | Introduced for pro and enterprise editions. |
*Note: When using the Async HTTP client you will have access to the request builder from ScrapeableFileEventData.getRedirectRequestBuilder() which can be used to modify and adjust the request before it is sent. If you use the Apache HTTP client the getRedirectRequestBuilder() method will always return null.
| Version | Description |
|---|---|
| 6.0.55a | Introduced for pro and enterprise editions. |
Returns the RedirectToURL value for the object.
This method does not receive any parameters.
Returns the RedirectToURL value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
| Version | Description |
|---|---|
| 6.0.55a | Introduced for pro and enterprise editions. |
*Note: Calling a setVariable or getVariable method in here WILL trigger the events for those again. Avoid infinite recursion please!
| Version | Description |
|---|---|
| 6.0.55a | Introduced for pro and enterprise editions. |
| Version | Description |
|---|---|
| 6.0.55a | Introduced for pro and enterprise editions. |
Creates an EventHandler callback object which will be called when the event triggers
| Version | Description |
|---|---|
| 6.0.55a | Introduced for pro and enterprise editions. |
Returns the name of the handler. This method doesn't need to be implemented but helps with debugging.
This method does not receive any parameters.
Returns the name of the handler. This method doesn't need to be implemented but helps with debugging.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Processes the event, and potentially returns a useful value modifying something in the internal code as defined by the EventFireTime used to launch this event.
Returns a value based on which AbstractEventData class is used.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
The AbstractEventData class is an abstract class which allows for the accessing of various data values found within ScreenScraper. Below are the various classes that extend AbstractEventData
AbstractEventData is extended by the following classes and it is those classes that should be used in place of AbstractEventData.
Returns the LastReturnValue for the object. This is the value previously returned by another callback. This can be null, if no callbacks have been fired yet for this event. A null value is also the default return value for the given event.
This method does not receive any parameters.
Returns the LastReturnValue for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Sets the LastReturnValue fro the object.
Returns void.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
ExtractorPatternEventData extends AbstractEventData
This contains the data for various extractor pattern operations
Inherits the following methods from AbstractEventData
Returns the status of the extractor pattern timeout. Returns true if and only if the extractor pattern was applied and timed out while doing so. Otherwise it will return false.
This method does not receive any parameters.
Returns a boolean value representing the status of the extractor pattern timeout.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the DataRecord value for the object.
This method does not receive any parameters.
Returns the DataRecord value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the DataSet value for the object.
This method does not receive any parameters.
Returns the DataSet value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the ExtractorPattern value for the object.
This method does not receive any parameters.
Returns the ExtractorPattern value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the Scrapeablefile value for the object.
This method does not receive any parameters.
Returns the Scrapeablefile value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the Session value for the object.
This method does not receive any parameters.
Returns the Session value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
ScrapeableFileEventData extends AbstractEventData
This contains the data for various scrapeable file operations
Inherits the following methods from AbstractEventData
Returns the HttpResponseData for the object.
This method does not receive any parameters.
Returns the HttpResponseData for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the RedirectRequestBuilder for the object. Use this to add headers, etc... for the redirect. It can be null depending on the HTTP client being used, and whether or not it supports manually playing with the redirect.
This method does not receive any parameters.
Returns the RedirectRequestBuilder for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the Scrapeablefile value for the object.
This method does not receive any parameters.
Returns the Scrapeablefile value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the Session value for the object.
This method does not receive any parameters.
Returns the Session value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
ScriptEventData extends AbstractEventData
This contains the data for various script operations
Inherits the following methods from AbstractEventData
Returns the DataRecord value for the object.
This method does not receive any parameters.
Returns the DataRecord value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the DataSet value for the object.
This method does not receive any parameters.
Returns the DataSet value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the Scrapeablefile value for the object.
This method does not receive any parameters.
Returns the Scrapeablefile value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the ScriptException for the object.
This method does not receive any parameters.
Returns the ScriptException for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the ScriptName value for the object.
This method does not receive any parameters.
Returns the ScriptName value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the Session value for the object.
This method does not receive any parameters.
Returns the Session value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
SessionEventData extends AbstractEventData
This contains the data for various session operations
Inherits the following methods from AbstractEventData
Returns the IncrementRecordsAmount value for the object.
This method does not receive any parameters.
Returns the IncrementRecordsAmount value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the Session value for the object.
This method does not receive any parameters.
Returns the Session value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the VariableName value for the object.
This method does not receive any parameters.
Returns the VariableName value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Returns the VariableValue value for the object.
This method does not receive any parameters.
Returns the VariableValue value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
StringEventData extends AbstractEventData
This contains the data for various string operations
Inherits the following methods from AbstractEventData
Returns the Input value for the object.
This method does not receive any parameters.
Returns the Input value for the object.
| Version | Description |
|---|---|
| 6.0.55a | Available for all editions. |
Add to the value of a session variable.
Returns void. If the variable doesn't exist, or is not a string or integer, a message will be added to the log. If it cannot add to the variable for any other reason it will write an error to the log.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Pause scrape and display breakpoint window. If the scrape is running in server mode, to avoid the break, logVariables will be called in place of breakpoint.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
Remove all session variables.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Clear stored cookies.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Clears the value of all session variables that match the keys in the Map. This will ignore a key of DATARECORD.
This method is provided using a Map or Collection rather than a List or Set to work easier with the setSessionVariables method.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Changed from session.removeSessionVariablesInMap to session.clearVariables. |
Decode HTML Entities on a session variable.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Downloads the file to the local file system.
Returns true on successful download of the file otherwise it return false.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. Lazy scrape only available for enterprise edition. |
If the file to download requires that POST data is sent in order to get the file you would use saveFileOnRequest with a scrapeable file.
Using this method in a script takes the place of requesting the target URL as a scrapeable file.
Manual start the execution of a script.
Returns void. If the file doesn't exist a message will be written to the log. If the called script has an error in it a warning will be written to the log.
| Version | Description |
|---|---|
| 5.0 | Scripts called using this method are now exported with the scraping session. |
| 4.5 | Available for professional and enterprise editions. |
Executes the named script, but preserves the current context (dataRecord, scrapeableFile, etc...)
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Get the general character set being used in page response renderings.
This method does not receive any parameters.
Returns the character set applied to the scraping session's files, as a string. If a character set has not been specified then it will default to the character set specified in settings dialog box.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
If you are having trouble with characters displaying incorrectly, we encourage you to read about how to go about finding a solution using one of our FAQs.
Retrieve the timeout value for scrapeable files in the session.
This method does not receive any parameters.
Returns the timeout value in milliseconds, as an integer.
| Version | Description |
|---|---|
| 5.0.1a | Introduced for all editions. |
Get the current cookies.
This method does not receive any parameters.
Returns an array of the cookies in the session.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Checks to see if this is currently set to run in debug mode. This is useful for developing scrapes, as enabling debug mode logs a warning message, so it is easier to notice a scrape with hard-coded values used for development. Also logs a warning in the web interface or log each time monitored variables are logged with the logMonitoredValues or webMessage methods are called.
This method takes no parameters.
True if debug mode is enabled, false otherwise.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Gets the default retry policy to be used by each scrapeable file when one wasn't set for it.
This method takes no parameters
The default return policy, or null if there isn't one
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Get how long the current session has been running.
This method does not receive any parameters.
Returns number of milliseconds the scrape has been running, as a long (8-byte integer).
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
If you would like to log the running time of the scraping session you should use logElapsedRunningTime.
Get the logging level of the scrape.
This method does not receive any parameters.
Returns the logging level, as an integer. Currently there are four levels: 1 = Debug, 2 = Info, 3 = Warn, 4 = Error.
| Version | Description |
|---|---|
| 5.0.1a | Introduced for all editions. |
Retrieve the maximum number of concurrent file downloads being allowed.
This methods does not receive any parameters.
Returns the max number of concurrent file downloads allowed, as an integer.
| Version | Description |
|---|---|
| 5.0 | Added for professional and enterprise editions. |
Retrieve the number of attempts that scrapeable files should make to get the requested page.
This method does not receive any parameters.
Returns the number of attempts that will be made, as a integer.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Get the total number of scripts allowed on the stack before the scraping session is forcibly stopped.
This method does not receive any parameters.
Returns max number of scripts that can be running at a time, as an integer.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Get the name of the current scraping session.
This method does not receive any parameters.
Returns the name of the scraping session, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Get the number of scripts currently running.
This method does not receive any parameters.
Returns number of running scripts, as an integer.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Determine whether or not non-tidied HTML is to be retained for all scrapeable files in this scraping session.
This method does not receive any parameters.
Returns whether non-tidied HTML is be retained for all scrapeable files or not, as a boolean.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Get the unique identifier for the scraping session.
This method does not receive any parameters.
Returns unique session id for the scraping session, as an integer.
| Version | Description |
|---|---|
| 5.0 | Added for enterprise edition. |
Retrieve the time at which the scrape started.
This method does not receive any parameters.
Returns the start time of the scrape in milliseconds, as a long.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Gets the current time zone of the Scraping Session
This method takes no parameters.
The time zone this scrape is set to.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Retrieve the value of a saved session variable.
Returns the value of the session variable. This will be a string unless you have used setVariable to place something other than a string into a session variable.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve the value of a saved session variable (alias of getVariable).
Returns the value of the session variable. This will be a string unless you have used setVariable to place something other than a string into a session variable.
| Version | Description |
|---|---|
| 4.5 | Added for all editions. |
Returns whether or not we are currently running in the command line. This is a convenience method for doing something different in a script when running in the command line as opposed to other modes
This method does not receive any parameters.
Returns true if and only if the scrape is currently running in the command line.
| Version | Description |
|---|---|
| 6.0.37a | Introduced for all editions. |
Returns whether or not we are currently running in the server. This is a convenience method for doing something different in a script when running in the server as opposed to other modes
This method does not receive any parameters.
Returns true if and only if the scrape is currently running in the server.
| Version | Description |
|---|---|
| 6.0.37a | Introduced for all editions. |
Returns whether or not we are currently running in the workbench. This is a convenience method for doing something different in a script when running in the workbench as opposed to other modes
This method does not receive any parameters.
Returns true if and only if the scrape is currently running in the workbench.
| Version | Description |
|---|---|
| 6.0.37a | Introduced for all editions. |
Loads the state that would have been previously saved by invoking the session.saveStateToString method.
None
| Version | Description |
|---|---|
| 5.5.30a | Available in Professional and Enterprise editions. |
Load session variables from a file.
Returns void. If there is a problem retrieving the file contents an I/O error will be written to the log.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
See also: saveVariables.
If you want to create your own file of session variables, the format is a hard return-delimited list of name/value pairs. Both the key and value should be URL-encoded.
Saves the current state of the scraping session to a string. An example use case for this method would be a scraping session that logs in to a site, extracts some information, and then is stopped, saving its state out to a file. A second scraping session could then be run, loading the state back in from the file, which would keep the session logged in so that other information could be obtained without logging in once again. By default the scraping session will save out information such as the URL to use as a referer. More information can be saved using the boolean flags described below.
None
| Version | Description |
|---|---|
| 5.5.30a | Available in Professional and Enterprise editions. |
Saves all current string and integer variables to a file.
Returns void. If there is a problem retrieving the file contents an I/O error will be written to the log.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Manually scrape a scrapeable file.
Returns void. If there is a problem accessing the scrapeable file an message will be written to the log.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Invokes a scrapeable file using a string of content instead of a web page or local file.
None
| Version | Description |
|---|---|
| 5.5.13a | Available in all editions. |
Send data to the external script that initiated the scrape. This isn't currently supported with all drivers (e.g., remote scraping session), check the documentation on the language of the external script for more information.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Set the general character set used in page response renderings. This can be particularly helpful when the pages render characters incorrectly.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
This method must be invoked before the session starts.
If you are having trouble with characters displaying incorrectly, we encourage you to ready about how to go about finding a solution using one of our FAQs.
Set the timeout value for scrapeable files in the session.
Returns void.
| Version | Description |
|---|---|
| 5.0.1a | Introduced for all editions. |
Manually set a cookie in the current session state.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for professional and enterprise editions. |
This method should be rarely used as screen-scraper automatically manages cookies. In cases where cookies are set via JavaScript, this function might be necessary.
Sets the debug state for the scrape. Enabled debug mode simply outputs a warning periodically while running, to help prevent running a production scrape in debug mode.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Sets a retry policy that will affect all files in the scrape. This policy will be used by all scrapeable files that do not have a retry policy set for them. If a retry policy was manually set for them, this one will not be used.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Sets the path to the keystore file. Some web sites require a special type of authentication that requires the use of a keystore file. See our blog entry on Using Client Certificates for more detail. Calling this method is the equivalent of setting the corresponding value under the "Advanced" tab for the scraping session in the workbench.
None
| Version | Description |
|---|---|
| 5.5.10a | Available in all editions. |
Sets the password for the keystore file. Some web sites require a special type of authentication that requires the use of a keystore file. See our blog entry on Using Client Certificates for more detail. Calling this method is the equivalent of setting the corresponding value under the "Advanced" tab for the scraping session in the workbench.
None
| Version | Description |
|---|---|
| 5.5.10a | Available in all editions. |
Set the logging level of the scrape.
Returns void.
| Version | Description |
|---|---|
| 5.0.1a | Introduced for all editions. |
Set the maximum number of concurrent file downloads to a allow.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for professional and enterprise editions. |
Set the number of attempts that scrapeable files should make to get the requested page.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
Get the total number of scripts that can be running concurrently. Default value for maxScriptsOnStack is 50.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for enterprise edition. |
Before you start upping the value of the number of scripts that can be on the stack you should make sure that your scrape is not eating more then it should. One thing to consider is recursion instead of iterating. This is discussed in more details on our blog or in the Tips, Tricks, and Samples section of this site.
Causes the "User-Agent" header sent by screen-scraper to be randomized. The user agent strings from which screen-scraper will select are found in the "resource\conf\user_agents.txt" file.
None
| Version | Description |
|---|---|
| 5.5.34a | Available in Professional and Enterprise editions. |
Set whether or not non-tidied HTML is to be retained for all scrapeable files.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
If, after the file is scraped, you want to be able to use getNonTidiedHTML this method has to be called before a file is scraped.
Sets the value of all session variables that match the keys in the Map to the values in the Map. This will ignore a key of DATARECORD.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
| 5.5.43a | Changed from session.setSessionVariablesFromMap to session.setSessionVariables. |
Sets a status message to be displayed in the web interface.
None
| Version | Description |
|---|---|
| 5.5.32a | Available in Enterprise edition. |
If this method is passed the value of true, it will cause screen-scraper to stop the current scraping session if an extractor pattern timeout occurs.
None
| Version | Description |
|---|---|
| 5.5.36a | Available in Professional and Enterprise editions. |
If this method is passed the value of true, it will cause screen-scraper to stop the current scraping session if the maximum attempts to request a file is reached.
None
| Version | Description |
|---|---|
| 5.5.36a | Available in Professional and Enterprise editions. |
If this method is passed the value of true, it will cause screen-scraper to stop the current scraping session if a script error occurs.
None
| Version | Description |
|---|---|
| 5.5.36a | Available in Professional and Enterprise editions. |
Sets the time zone that will be used when using a method that returns a time formatted as a string.
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
If this method is passed the value of true, it will cause screen-scraper to utilize whatever character set is specified by the server in its "Content-Type" response header. If no such header exists, screen-scraper will default to either the character set indicated for the scraping session or the global character set (in that order).
None
| Version | Description |
|---|---|
| 5.5.11a | Available in all editions. |
Sets the user agent to be used for all requests.
None
| Version | Description |
|---|---|
| 5.5.23a | Available in Professional and Enterprise editions. |
Set the value of a session variable.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Set the value of a session variable (alias of setVariable).
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Determine if the scrape has been stopped. This can be done using the stop button in the workbench or the stop scraping button on the web interface (for enterprise users).
This method does not receive any parameters.
Returns true if the scrape has been requested to stop; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for enterprise edition. |
Stop the current scraping session.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Waits for any file downloads to complete before returning. This should be used in tandem with the session.downloadFile method call that takes the "doLazy" paraameter.
None
None
| Version | Description |
|---|---|
| 5.5.43a | Available in Enterprise edition. |
The sutil class provides general functions used to manipulate and work with extracted data. It also allows you to get information regarding screen-scraper such as its memory usage or version.
In the course of a scrape it you might want to gather images associated with the other information being gathered. These methods are provided to not only download the images but to gather size information and resize to your desired size.
These methods are only available to enterprise edition users.
Get the height of an image.
Returns the height in pixels of the image file, as an integer. If the file doesn't exist or is not an image an error will be thrown and -1 will be returned.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for enterprise edition. |
Get the width of an image.
Returns the width in pixels of the image file, as an integer. If the file doesn't exist or is not an image an error will be thrown and -1 will be returned.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for enterprise edition. |
Internally, only one function is used to resize all images; however, to facilitate the resizing of images, we have provided you with three methods. Each method will help you specify what measurement is most important (width or height) and whether the image should retain its aspect ratio.
Resize image, retaining aspect ratio, based on specified height.
Returns void. If an error is encountered it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for enterprise edition. |
Resize image, retaining aspect ratio, based on specified width.
Returns void. If an error is encountered it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for enterprise edition. |
Resize image to a specified size.
Returns void. If an error is encountered it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for enterprise edition. |
This method can cause distortions of the image if the aspect ratio of the original and target images are different.
To be used in conjunction with the ImageDecoder class.
This class represents decoded images. The objects can be queried for the text that was in the image, as well as any error that occurred while the image was being decoded. When the returned text is incorrect, there is a method that can be used to report it as bad. This can be used for sites like decaptcher.com, where refunds are given for incorrectly interpreted images.
Gets any error message, or returns null if there was no error
This method takes no parameters
The error message returned
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Gets the result from decoding the image. Most likely this will be a String, but each implementation could return a specific object type.
This method takes no parameters
The text extracted from the image, or null if there was an error
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Handles an incorrectly resolved image. Some types of decoders won't have anything here
This method takes no parameters
This method returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Returns true if there was an error, false otherwise. Also returns false if the image has not been resolved yet
This method takes no parameters
True if there was an error, false otherwise
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Class to convert images to text for interacting with CAPTCHA challenges. There are currently two implementations:
When a reference to an image is passed to an instance of this class, it returns a DecodedImage object that can be queried for the resulting text, errors, and can report an image as poorly converted.
See example attached.
Requires an account with decaptcher.com.
Type of ImageDecoder in the com.screenscraper.util.images package that uses the decaptcher.com service to convert images to text. The constructor is DecaptcherDecoder(ScrapingSession session, String username, String password) or DecaptcherDecoder(ScrapingSession session, String username, String password, String apiUrl).
Returns void. If it runs into any problems accessing the decaptcher.com service an error will be thrown.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions |
| 5.5.40a | Added the port parameter. The service now requires the correct port in order to authenticate. |
Initialization script
Type of ImageDecoder in the com.screenscraper.util.images package that uses a popup window prompting the user to enter the text read from an image. Useful for debugging purposes, as the input text should always be correct (so long as it is typed correctly). Helpful during testing to avoid costs associated with paid-for CAPTCHA decoding services such as decaptcher.com.
Returns void. If it runs into any problems decoding an image an error will be thrown.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions |
Initialize script
Converts the image given to a DecodedImage that will handle it. Does not delete the file.
A DecodedImage used to get the text, errors, and possibly report a result as bad.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Converts the image at the given URL to a DecodedImage that will handle it. Temporarily saves the file in the screen-scraper root folder, but deletes it once it has been decoded. By default, this will use the scraping session's HttpClient to request the URL.
A DecodedImage used to get the text, errors, and possibly report a result as bad.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Converts the Date given to a string in a specified format, or in the "MM/dd/yyyy HH:mm:ss.SS zzz" if no format is given.
A String representing the date given
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Decode HTML Entities.
Returns string with decoded HTML entities.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Converts a String to a Date object using the given format. If null is given as a format, "MM/dd/yyyy HH:mm:ss.SS zzz" is used
The Date object matching the date given in the String, or null if it couldn't be parsed with the given format
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Replaces the UTF variants on whitespace with a regular space character.
Returns the converted string.
| Version | Description |
|---|---|
| 6.0.55a | Available in all editions. |
Checks to see if one date is within a certain number of days of another.
| Version | Description |
|---|---|
| 5.5.13a | Available in all editions. |
Compare two strings ignoring case.
Returns true if the values of the two strings are equal when case is not considered; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Returns a number formatted in such a way that it could be parsed as a Float, such as xxxxxxxxx.xxxx. It attempts to figure out if the number is formatted as European or American style, but if it cannot determine which it is, it defaults to American. If the number is something with a k on the end, it will convert the k to thousand (as 000). It will also try to convert m for million and b for billion. It also assumes that you won't have a number like 3.123k or 3.765m, however 3.54m is fine. It figures if you wanted all three of those digits you would have specified it as 3765k or 3,765k
Returns a String formatted as a phone number, such as +1 (123) 456-7890x2, or null if the input was null
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Converts a String to a US formatted phone number, as +1 (123) 456-7890x2. Expects a 7 digit or 10+ digit phone number. The extension is optional, and will be any digits found after an x. This allows for extensions listed as ext, x, or extension.
Returns a String formatted as a phone number, such as +1 (123) 456-7890x2, or null if the input was null
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Formats and returns a US style zip code as 12345-6789. If the given zip code isn't 5 or 9 digits, will log a warning, but it will put 5 digits before the - and anything else (if any) after the -
Zip code formatted String, such as 12345-6789 or 12345
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Returns the current date in a specified format, or uses the "MM/dd/yyyy HH:mm:ss.SS zzz" if null is given. Uses the session's timezone.
A String representing the date and time this method was invoked
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Retrieve the file path of the screen-scraper installation.
This method does not receive parameters.
Returns the installation directory file path, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Get memory usage of screen-scraper.
This method does not receive any parameters.
Returns the average percentage of memory used by screen-scraper over the past 30 seconds, as an integer.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for enterprise edition. |
For tips on optimizing screen-scraper's memory usage so that it can run faster, see our FAQ on optimization.
Get the mime-type of a local file.
Returns the mime-type of the file, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Get the number of runnable scraping sessions.
This method does not receive any parameters.
Returns the number of scraping sessions in this instance of screen-scraper, as a integer.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Gets the number of scraping sessions that are currently being run.
An int representing the number of running scraping sessions.
| Version | Description |
|---|---|
| 5.5.42a | Available in Enterprise edition. |
Gets a DataSet containing each of the elements of a <select> tag. The returned DataRecords will contain a key for the text found between the tags (possibly with html tags removed), a value indicating if it was the selected option, and the value to submit for the specific option. Note that this only looks for option tags, and as such passing in text containing more than a single select tag will produce false output.
A DataSet with one record per option. Values extracted will be stored in
VALUE : The value the browser would submit for this option
TEXT : The text that was between the tags
SELECTED : A boolean that is true if this option was selected by default
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Gets all the options from a radio button group. The values are returned in a data record. Any labels that are to be ignored will not be included in the returned set. Not all buttons have a label, as radio buttons do not require a label, and it would be difficult to know in a regular expression exactly what to extract as the label unless there is a label tag.
DataSet containing one record for each of the extracted radio buttons. Values will be stored in
VALUE : The value the browser would submit for this radio button
TEXT : The text that represents this button, or null if no label could be found for it
SELECTED : A boolean that is true if this button was selected by default
ID : The ID of the radio button, or null if no ID was found
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Gets a random referrer page from a list of many different search engine web sites and a few other pages.
This method does not receive any parameters.
Returns a random referrer URL.
| Version | Description |
|---|---|
| 6.0.1a | Introduced for all editions. |
Returns a random User Agent. The list isn't closely monitored, so it may not include newer user agents, and may include extremely old ones as well.
This method does not receive any parameters.
Returns a random user agent.
| Version | Description |
|---|---|
| 6.0.1a | Introduced for all editions. |
Get edition of screen-scraper instance.
This method does not receive any parameters.
Returns the edition name, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Get version of screen-scraper instance.
This method does not receive any parameters.
Returns the version number, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Determine if the value of a string is an integer.
Returns true if the string is an integer; otherwise, it returns false. If it is passed an object that is not a string, including an integer, an error will be thrown.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Determine if an object's value is null or empty.
Returns true if the value of the object is null or an empty string; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Determine if operating system is a Linux platform.
This method does not receive parameters.
Returns true if the operating system is Linux; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Determine if operating system is a Mac platform.
This method does not receive parameters.
Returns true if the operating system is Mac; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Determine if operating system is a Windows platform.
This method does not receive parameters.
Returns true if the operating system is Windows; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Retrieve the response contents of a GET request.
Returns contents of the response, as a string.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
This method will use any proxy settings that have been specified in the Settings dialog box.
Makes a GET request and returns the result as a string. This method will use the proxy settings indicated in the "Settings" dialog box, if any.
This method does not receive any parameters.
| Version | Description |
|---|---|
| 6.0.6a | Introduced for all editions. |
Makes a GET request and returns the result as a string. This method will use the proxy settings attached to the current scraping session.
This method does not receive any parameters.
| Version | Description |
|---|---|
| 6.0.6a | Introduced for all editions. |
Retrieve the response header contents.
Returns contents of the response, as a two-dimensional array.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
This method will use any proxy settings that have been specified in the Settings dialog box..
Merges two data records by copying all values from the second record over values of the first record, and returning a new DataRecord with these values. Doesn't modify either original record
A new DataRecord with the merged values
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Get an object in string format.
Returns an empty string if the value of the object is null; otherwise, returns the value of the toString method of the object.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Attempts to parse a string to a name. The parser is not perfect and works best on english formatted names (for example, "John Smith Jr." or "Guerrero, Antonio K". This uses standard settings for the parser. To get more control over how the name is parsed, use the EnglishNameParser class.
Returns the parsed name, as a Name object.
| Version | Description |
|---|---|
| 6.0.59a | Available for professional and enterprise editions. |
Attempts to parse a string to a name. The parser is not perfect and works best on english formatted names (for example, "John Smith Jr." or "Guerrero, Antonio K". This uses standard settings for the parser. To get more control over how the name is parsed, use the EnglishNameParser class.
Returns the parsed name, as a Name object.
| Version | Description |
|---|---|
| 6.0.59a | Available for professional and enterprise editions. |
Attempts to parse a string to an address. The parser is not perfect and works best on US addresses. Most likely other address formats can be parsed with the USAddressParser class by providing different constraints in the builder. This method is here for convenience in working with US addresses.
Returns the parsed address, as a Address object.
| Version | Description |
|---|---|
| 6.0.59a | Available for professional and enterprise editions. |
Pause scraping session.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for professional and enterprise editions. |
Pausing the scraping session also pauses the execution of the scripts including the one that initiates the pause.
Pauses for a random amount of time. This is also setup to stop immediately if the stop scrape button is clicked, and to allow breakpoints to be triggered while it is pausing.
Returns void.
| Version | Description |
|---|---|
| 5.5.29a | Available in professional and enterprise editions. |
Change a date format.
Returns formatted date according to the specified format, as a string.
| Version | Description |
|---|---|
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for professional and enterprise editions. Unspecified source format available for enterprise edition. |
The date formats are not the same for the two methods. Read carefully.
Send an email using SMTP mail server specified in the settings.
Returns void. If it runs into any problems while attempting to send the email an error will be thrown.
| Version | Description |
|---|---|
| 6.0.35a | Now supports alternate content types. |
| 5.0 | Moved from session to sutil. |
| 4.5 | Available for enterprise edition. |
Sorts the elements in a set into an ordered list.
This method returns a sorted list of elements that are in the set.
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
Determine if one string is the start of another, without regards for case.
Returns true if string starts with start when case is not considered; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Parse string into a floating point number.
Returns the string's value as a floating point number.
| Version | Description |
|---|---|
| 5.0.1a | Introduced for professional and enterprise editions. |
Strips HTML from a string, replacing some tags with corresonding text-only equivalents.
Returns the stripped content.
| Version | Description |
|---|---|
| 6.0.20a | Available in only the Enterprise edition. |
Tidies the DataRecord by performing actions based on the values of the settings map given (or getDefaultTidySettings if none is given). Each value in the record that is a string will be tidied. Keys are not modified. The record given will not be modified, but a new record with the tidied values will be returned. If no settings are given, will use the values obtained from sUtil.getDefaultTidySettings().
The settings tidy settings and their default values are given below. If a key is missing in the settings map, that operation will not be performed.
| Map Key | Default Value | Description of operation performed |
|---|---|---|
| trim | true | Trims whitespace from values |
| convertNullStringToLiteral | true | Converts the string 'null' (without quotes) to the null literal (unless it has quotes around it, such as "null") |
| convertLinks | true | Preserves links by converting <a href="link">text</a> to text (link), will try to resolve urls if scrapeableFile isn't null. Note that if there isn't a start and end <a> tag, this will do nothing |
| removeTags | true | Remove html tags, and attempts to convert line break HTML tags such as <br> to a new line in the result |
| removeSurroundingQuotes | true | Remove quotes from values surrounded by them -- "value" becomes value |
| convertEntities (professional and enterprise editions only) | true | Convert html entities |
| removeNewLines | false | Remove all new lines from the text. Replaces them with a space |
| removeMultipleSpaces | true | Convert multiple spaces to a single space, and preserve new lines |
| convertBlankToNull | false | Convert blank strings to null literal |
A new DataRecord containing all the tidied values and any values that were not Strings in the original record. The values that were Strings but were not tidied as well as the DATARECORD value will not be in the returned record.
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
| 5.5.28a | Now uses a Map for the settings, rather than bit flags. |
Tidies the string by performing actions based on the values of the settings map.
The tidy settings and their default values are given below. If a key is missing in the settings map, that operation will not be performed.
| Map Key | Default Value | Description of operation performed |
|---|---|---|
| trim | true | Trims whitespace from values |
| convertNullStringToLiteral | true | Converts the string 'null' (without quotes) to the null literal (unless it has quotes around it, such as "null") |
| convertLinks | true | Preserves links by converting <a href="link">text</a> to text (link), will try to resolve urls if scrapeableFile isn't null. Note that if there isn't a start and end <a> tag, this will do nothing |
| removeTags | true | Remove html tags, and attempts to convert line break HTML tags such as <br> to a new line in the result |
| removeSurroundingQuotes | true | Remove quotes from values surrounded by them -- "value" becomes value |
| convertEntities (professional and enterprise editions only) | true | Convert html entities |
| removeNewLines | false | Remove all new lines from the text. Replaces them with a space |
| removeMultipleSpaces | true | Convert multiple spaces to a single space, and preserve new lines |
| convertBlankToNull | false | Convert blank strings to null literal |
The tidied string
| Version | Description |
|---|---|
| 5.5.26a | Available in all editions. |
| 5.5.28a | Now uses a Map for the settings, rather than bit flags. |
Assuming the extracted text's HTML code was:
<a href="http://www.somelink.com">This</a> was great because of these reasons:<br />
1 - Some reason<br />
2 - Another reason<br />
3 - Final reason
The output text would be:
This (http://www.somelink.com) was great because of these reasons:
1 - Some reason
2 - Another reason
3 - Final reason
Unzip a zipped file. Contents will appear in the same directory as the zipped file.
Returns void. If a file input/output error is experienced it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
Write to a file.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Added for all editions. |
screen-scraper provides three built-in objects for proxy sessions. These objects are: proxySession, request, and response. See the Variable scope section for details on which objects are available based on when scripts are run.
This object gives you the ability to control interactions with the proxy session. It is only for use in scripts that associated with proxy sessions.
Retrieve the value of the proxy session variable.
Returns the value of the session variable.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Write to the log.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Set the value of a proxy session variable.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
A request objects references a proxySession page request. Through this object you can control various aspects of the request.
Scripts run in the scraping engine use the scrapeable file to manipulate server requests.
Manually add an HTTP header.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Add POST parameter to HTTP request.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve the URL of the request.
This method does not receive any parameters.
Returns the URL of the request, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Manually remove an HTTP header. Both the key and value have to be specified as HTTP headers allow for multiple headers with the same key.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Remove POST parameter from HTTP request.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Manually set the request line.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
The response class provides you with a means for editing the responses received by the proxy server.
Scripts run in the scraping engine us the scrapeable file to manipulate server responses.
Add HTTP header to response.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve the content of the response.
This method does not receive any parameters.
Returns the content of the response, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve the status line of the response.
This method does not receive any parameters.
Returns the status line of the response, as a string.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Remove HTTP header from response.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Manually set the response content.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Manually set the status line.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
There are many classes that can be very helpful in getting your scripts to run correctly. Many of these are initially developed in-house to speed up coding time and once they have proved very stable offered to the public. For all classes you will need to import their packages. They are not automatically imported like the built-in screen-scraper objects.
The Apache Lang library provides enhancements to the standard Lang library of Java and can be particularly useful for completing tasks. As it is not a class that we maintain we will not document the methods in case they change without our notice but we invite you to look over how to use it in their API.
The CSVReader is not a class that is part of screen-scraper but is very useful and well put together. We have used it extensively. It is part of the opencsv package which actually holds the under pinnings of our own CsvWriter. As it is not a class that we maintain we will not document the methods in case they change without our notice but we invite you to look over how to use it in their API or brief documentation.
To use the CSVReader simply import it in your script, the same as you would any other utility class. The opencsv.jar file is already included in the Professional and Enterprise Editions of screen-scraper's default installation.
This CsvWriter has been created to work particularly well with the screen-scraper objects. It is simple to use and provided to ease the task of keeping track of everything when creating a csv file.
The most used methods are documented here but if you would like more information you can read the JavaDoc for the CsvWriter.
Create a csv file writer.
Returns a CsvWriter object. If it encounters an error it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Available for Professional and Enterprise editions. |
| 4.5.18a | Introduced in alpha version. |
com.screenscraper.csv.CsvWriter
Clear the buffer contents and close the file.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
| 4.5.18a | Introduced in alpha version. |
Write the buffer contents to the file.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
| 4.5.18a | Introduced in alpha version. |
Set the header row of the csv document. If the document already exists the headers will not be written. Also creates a data record mapping to ease writing to file.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
| 4.5.18a | Introduced in alpha version. |
If you want to use the data record mapping then the extractor tokens names should be all caps and all spaces should be replaced with underscores.
Write to the CsvWriter object.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for all editions. |
| 4.5.18a | Introduced in alpha version. |
This class is used to instantiate a data manager object. This is done to simplify the process of creating a data manager of a given type. Currently it only creates SqlDataManagers. A SQL data manager can be created without the use of this class, but it is simplified greatly through its use.
This class should no longer be used. Use a java.sql.BasicDataSource or com.screenscraper.datamanager.SshDataSource instead. See the SqlDataManager.buildSchemas page for examples
This class is only available for Professional and Enterprise editions of screen-scraper.
This method is no longer supported. Use a java.sql.BasicDataSource or com.screenscraper.datamanager.SshDataSource instead. See the SqlDataManager.buildSchemas page for examples.
Create a MsSQL data manager object.
Returns a SqlDataManager object. If an error is experienced it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
In order to create the MsSQL data manager you will need to make sure to install the appropriate jdbc driver. This can be done by downloading the MsSQL JDBC driver and placing it in the lib/ext folder in the screen-scraper installation directory.
This method is no longer supported. Use a java.sql.BasicDataSource or com.screenscraper.datamanager.SshDataSource instead. See the SqlDataManager.buildSchemas page for examples.
Create a MySQL data manager object.
Returns a SqlDataManager object. If an error is experienced it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
In order to create the MySQL data manager you will need to make sure to install the appropriate jdbc driver. This can be done by downloading the MySQL JDBC driver and placing it in the lib/ext folder in the screen-scraper installation directory.
This method is no longer supported. Use a java.sql.BasicDataSource or com.screenscraper.datamanager.SshDataSource instead. See the SqlDataManager.buildSchemas page for examples.
Create an Oracle data manager object.
Returns a SqlDataManager object. If an error is experienced it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
In order to create the Oracle data manager you will need to make sure to install the appropriate jdbc driver. This can be done by downloading the Oracle JDBC driver and placing it in the lib/ext folder in the screen-scraper installation directory.
This method is no longer supported. Use a java.sql.BasicDataSource or com.screenscraper.datamanager.SshDataSource instead. See the SqlDataManager.buildSchemas page for examples.
Create a Postgre data manager object.
Returns a SqlDataManager object. If an error is experienced it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
In order to create the Postgre data manager you will need to make sure to install the appropriate jdbc driver. This can be done by downloading the Postgre JDBC driver and placing it in the lib/ext folder in the screen-scraper installation directory.
This method is no longer supported. Use a java.sql.BasicDataSource or com.screenscraper.datamanager.SshDataSource instead. See the SqlDataManager.buildSchemas page for examples.
Create a SQLite data manager object.
Returns a SqlDataManager object. If an error is experienced it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
In order to create the Sqlite data manager you will need to make sure to install the appropriate jdbc driver. This can be done by downloading the Sqlite JDBC driver and placing it in the lib/ext folder in the screen-scraper installation directory.
The proxy server pool object is used to aid with manual anonymization of scrapes. An example of how to setup manual proxy pools is available in the documentation. You will likely want to read that page first if you are new to the process.
Additionally, you should reference the available method's available in the Anonymous API
Initiate a ProxyServerPool object.
This method does not receive any parameters.
Returns a ProxyServerPool.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
com.screenscraper.util.ProxyServerPool
Set the timeout that will render a proxy as being bad.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retrieve the number of available proxy servers.
This method does not receive any parameters.
Returns the number of available proxy servers, as an integer.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Write list of proxies to log.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Add proxy servers to pool using a text file.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Enables or disables automatic proxy cycling. When this is set to false (default is true) the current proxy that was automatically selected from the pool will be used each time the next proxy is requested. When set to true, each call to the getNextProxy method will cycle as normal between all available proxies.
A boolean value.
None
| Version | Description |
|---|---|
| 5.5.17a | Available in Professional and Enterprise editions. |
Set the number of proxies that can be tested concurrently.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Set threshold to get more proxy servers.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Write list of proxies after invalid proxies have been removed.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for all editions. |
Retry Policies are objects that tell a scrapeable file how to check for errors, and optionally what to do before retrying to download the files. Some of the things that can be done are executing scripts when the page loads incorrectly or running Runnables. Usually these things would either request a new proxy, output some helpful information, or could simply stop the scrape. RetryPolicy is an interface and can be implemented to create a custom retry policy, or there is a RetryPolicyFactory class that can be used to create some standard policies.
This policy is checked AFTER all the extractors have been run. This allows for checks on whether extractor patterns matched or not, and also allows a page to have it's 'error status' based off of another page (since extractor patterns could execute scripts that scrape other files, and those files could set a variable that acts as a flag to a previous retry policy). This could also cause some problems if the scrape isn't built to handle a page whose extractors shouldn't be run before the error checking occurs.
This interface is in the com.screenscraper.util.retry package.
If you need a custom retry policy, you can implement your own version of it. Be aware that you will need to ensure the references it has to the scrapeableFile are to the correct scrapeableFile. This could be tricky if you use the session.setDefaultRetryPolicy method. When using the scrapeableFile.setRetryPolicy method, the scrapeableFile will be the correct object. The interface is given below.
To help ensure you can create custom retry policies that have access to the scraping session and the scrapeable file that is currently being checked, there is an AbstractRetryPolicy class in the same package as the interface. This class defines some default behavior and adds protected fields for the session and scrapeable file that get set before the policy is run. If you extend this abstract class you can access the session and scrapeable file through this.scrapingSession and this.theScrapeableFile. Due to some oddities with the interpreter it is best to reference these variables with 'this.' to eliminate a few problems that arise in a few specific cases.
Returns a map that can be used to output an error message to indicate what checks failed. For instance, you could set a key to the value "Status Code" and the value '200', or a key with "Valid Page" and value 'false'
This method takes no parameters
Map of keys, or null if no values are indicated
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Return the maximum number of times this policy allows for a retry before terminating in an error
This method takes no parameters
The maximum number of times to allow the ScrapeableFile to be rescraped before resulting in an error
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Checks to see if the page loaded incorrectly
This method takes no parameters
True on errors, false otherwise
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Returns true if the referrer should be reset before attempting to rescrape the file, if there was an error. This can be useful to reset so the referrer doesn't show the page you just requested.
This method takes no parameters
True if the referrer should be reset if there was an error, false otherwise.
| Version | Description |
|---|---|
| 6.0.36a | Available in all editions. |
Returns true if the session variables should be reset before attempting to rescrape the file, if there was an error. This can be useful especially if extractors null session variables when they don't match, but the value is needed to rescrape the file.
This method takes no parameters
True if session variables should be reset if there was an error, false otherwise.
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
This will be called if all the retry attempts for the scrapeable file failed. In other words, if the policy said to retry 25 times, after 25 failures this method will be called. Note that runOnError will be called just before this, as it is called after each time the scrapeable file fails to load correctly, including the last time it fails to load.
This should only contain code that handles the final error. Any proxy rotating, cookie clearing, etc... should generally be done in the runOnError method, especially since it will still be called after the final error.This method takes no parameters
This method returns void
| Version | Description |
|---|---|
| 6.0.37a | Available in all editions. |
Runs this code when the page had an error. This could include things such as rotating the proxy. This code will be executed just before the page is downloaded again.
This method takes no parameters
This method returns void
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Returns true if errors should be logged to the log/web interface when they occur
This method takes no parameters
True if errors should be logged to the log/web interface when they occur
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Class used to create simple Retry Policies. See the RetryPolicy page for more details on what a RetryPolicy does. This class is found in the com.screenscraper.util.retry package.
Policy that retries if there was an error on the request by status code. Executes the runnable given before retrying.
The RetryPolicy to set in the ScrapeableFile
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Policy that returns no error. Useful for having a session-wide retry policy, but then using this for a particular scrapeable file so it doesn't use the session's policy
The RetryPolicy to set in the ScrapeableFile
| Version | Description |
|---|---|
| 6.0.25a | Available in all editions. |
Policy that requires a Regular Expression to match the page content (including headers) in order to be considered valid.
The RetryPolicy to set in the ScrapeableFile
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
Policy that requires a Regular Expression NOT to match the page content (including headers) in order to be considered valid. In other words, if the Regular Expression matches, it means that the page should be rescraped.
The RetryPolicy to set in the ScrapeableFile
| Version | Description |
|---|---|
| 5.5.29a | Available in all editions. |
This object simplifies your interactions with a JDBC-compliant SQL database. It can work with various types of databases and even in a multi-threaded format to allow scrapes to continue without having to wait for the queries to process. View an example of how to use the SqlDataManager.
This feature is only available for Professional and Enterprise editions of screen-scraper.
Prefer a more traditional approach? See an example of Working with MySQL databases.
In order to use the SqlDataManager you will need to make sure to install the appropriate JDBC driver. This can be done by downloading the driver and placing it in the lib/ext folder in the screen-scraper installation directory.
Add an event callback to SqlDataManager object.
This feature is only available for Professional and Enterprise editions of screen-scraper.
Before adding an event to the SqlDataManager, you must build the schema of any tables you will use because events are related to table operations such as inserting data
public void handleEvent(DataManagerEvent event) that needs to be implemented. The DataManagerEvent has a method getDataNode() to retrieve the relevant DataNode.Returns a DataManagerEventListener. The same DataManagerEventListener object that was passed in
| Version | Description |
|---|---|
| 5.5 | Available for professional and enterprise editions. |
Add data to fields, in preparation for insertion into a database.
When adding data in a many-to-many relation, if setAutoManyToMany is set to false, a null row should be inserted into the relating table so the datamanager will link the keys correctly between related tables. For example, dm.addData("many_to_many", null);
Before adding data the first time, you must build the schema of any tables you will use, as well as add foreign keys if you are not using a database engine that natively supports them (such as InnoDB for MySQL).
The SqlDataManager will attempt to convert a value that is given to the correct format for the database. For example, if the database requires an int for a column named age, dm.addData("table", "age", "32") will convert the String "32" to an int before adding it to the database. See the table below the examples for other types of java objects and how they map to SQL types.
The table and columnName parameters are not case sensitive. The same is true for the key values in the data map.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Since the DataManager is designed with screen-scraper in mind all inputs support using the String type in addition to their corresponding Java object type, but the String needs to be parseable into the corresponding data type. For example if there is a column that is defined as an Integer in the database then the String needs to be parseable by Integer.parseInt(String value). Here is a mapping of the sql types (based on java.sql.Types) to Java objects:
| SQL Type | Java Object | |
|---|---|---|
| java.sql.Types.CHAR | String | |
| java.sql.Types.VARCHAR | String | |
| java.sql.Types.LONGVARCHAR | String | |
| java.sql.Types.LONGNVARCHAR | String | |
| java.sql.Types.NUMERIC | BigDecimal | |
| java.sql.Types.DECIMAL | BigDecimal | |
| java.sql.Types.TINYINT | Integer | |
| java.sql.Types.SMALLINT | Integer | |
| java.sql.Types.INTEGER | Integer | |
| java.sql.Types.BIGINT | Long | |
| java.sql.Types.REAL | Float | |
| java.sql.Types.FLOAT | Double | |
| java.sql.Types.DOUBLE | Double | |
| java.sql.Types.BIT | Boolean | |
| java.sql.Types.BINARY | ByteArray | |
| java.sql.Types.VARBINARY | ByteArray | |
| java.sql.Types.LONGVARBINARY | ByteArray | |
| java.sql.Types.DATE | SQLDate or Long | |
| java.sql.Types.TIME | SQLTime or Long | |
| java.sql.Types.TIMESTAMP | SQLTime or Long | |
| java.sql.Types.ARRAY | Object | |
| java.sql.Types.BLOB | ByteArray | |
| java.sql.Types.CLOB | Object | |
| java.sql.Types.JAVA_OBJECT | Object | |
| java.sql.Types.OTHER | Object |
Manually setup table connection (key matching).
If SqlDataManager.buildSchemas is called, any foreign keys manually added before that point will be overridden or erased.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
If the database has some indication of foreign keys then these will be followed automatically. If the database does not allow for foreign key references then you will need to build the table connections using this method.
Manually add session variable data to fields, in preparation for insertion into a database.
The keys from the session will be matched in a case insensitive way to the column names of the database.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Add corresponding session variables to the tables automatically when it is committed.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Collect the database schema information, including foreign key relations between tables.
Schemas must be built for any tables that will be used by this DataManager before data can be added.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Clear all data from the data manager without writing it to the database. This includes all data previously committed but not yet written.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Clear session variables corresponding to the fields of a specific table (case insensitive).
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Clear session variables corresponding to a committed table automatically.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Close data manager's connections.
If there is data that has not yet been written to the database when this method is called it will not be written.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Commit a prepared row of data into queue. Once called the data can no longer be edited. When working with multiple tables that relate by a foreign key, it is important to commit rows in the correct order. The rows in each of the child tables should be committed before the parent, or they will not be correctly linked when written to the database.
This does not write the row of data to the database, but rather puts it in queue to be written at a later time.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Commit prepared rows of data for all tables into queue. Once called the data can no longer be edited.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Write committed data to the database. Any data that has not been committed using either the commit or commitAll method will be lost and not written to the database.
This method does not receive any parameters.
Returns true data was successfully written to the database; otherwise, it returns false.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Retrieve the connection object of the data manager. This can be helpful if you want to do something that the data manager cannot do easily, such as query the database.
Be sure to close the connection once it is no longer needed. Failure to do so could exhaust the connection pool used by the datamanger, which will cause the scraping session to hang.
This method does not receive and parameters.
Returns a connection object matching the one used in the data manager.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Retrieve the last autogenerated primary key, if any, for the given table
case insensitve table name
Returns a com.screenscraper.datamanager.DataObject containing the primary key.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Sets whether or not the data manager should automatically take care of many-to-many relationships.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
If the many-to-many table has more information than just the keys then you will want to leave this feature turned off so that you can add more data than just the keys before committing.
This feature is only available for Professional and Enterprise editions of screen-scraper.
Set global merge status. When conflicts exist in data, a merge of true will take the newer values and save them over previous null values.
When merging or updating values in a table, that table must have a Primary Key. When the Primary Key is set to autoincrement, if the value of that key was not set with the addData method the DataManager will create a new row rather than update or merge with an existing row. One solution is to use an SqlDuplicateFilter to set fields that would identify an entry as a duplicate and automatically insert the value of the autoincrement key when data is committed.
| Update | Merge | Resulting Action |
|---|---|---|
| false | false | Ignore row on duplicate |
| true | false | Update only values whose corresponding columns are currently NOT NULL in the database |
| false | true | Update only values whose corresponding columns are currently NULL in the database |
| true | true | Update all values to new data |
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
This feature is only available for Professional and Enterprise editions of screen-scraper.
Set update status globally. When conflicts exist in data, an update of true will take the newer values and save them over previous non-null values.
When merging or updating values in a table, that table must have a Primary Key. When the Primary Key is set to autoincrement, if the value of that key was not set with the addData method the DataManager will create a new row rather than update or merge with an existing row. One solution is to use an SqlDuplicateFilter to set fields that would identify an entry as a duplicate and automatically insert the value of the autoincrement key when data is committed.
| Update | Merge | Resulting Action |
|---|---|---|
| false | false | Ignore row on duplicate |
| true | false | Update only values whose corresponding columns are currently NOT NULL in the database |
| false | true | Update only values whose corresponding columns are currently NULL in the database |
| true | true | Update all values to new data |
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Set the error logging level. Currently only DEBUG and ERROR levels are supported. At the DEBUG level, all queries and results will be output to the log.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
This feature is only available for Professional and Enterprise editions of screen-scraper.
Set merge status for a table. When conflicts exists in data, a merge of true will take the newer values and save them over previous null values.
When merging or updating values in a table, that table must have a Primary Key. When the Primary Key is set to autoincrement, if the value of that key was not set with the addData method the DataManager will create a new row rather than update or merge with an existing row. One solution is to use an SqlDuplicateFilter to set fields that would identify an entry as a duplicate and automatically insert the value of the autoincrement key when data is committed.
| Update | Merge | Resulting Action |
|---|---|---|
| false | false | Ignore row on duplicate |
| true | false | Update only values whose corresponding columns are currently NOT NULL in the database |
| false | true | Update only values whose corresponding columns are currently NULL in the database |
| true | true | Update all values to new data |
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Set number of threads that the data manager can have open at once. When set higher than one, the scraping session can continue to run and download pages while the database is being written. This can decrease the time required to run a scrape, but also makes debugging harder as there is no guarantee about the order in which data will be written. It is recommended to leave this setting alone while developing a scrape. Also, the flush method will always return true if more than one thread is being used to write to the database, even if the write failed.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
This feature is only available for Professional and Enterprise editions of screen-scraper.
Set update status for a given table. When conflicts exists in data, an update of true will take the newer values and save them over previous non-null values.
When merging or updating values in a table, that table must have a Primary Key. When the Primary Key is set to autoincrement, if the value of that key was not set with the addData method the DataManager will create a new row rather than update or merge with an existing row. One solution is to use an SqlDuplicateFilter to set fields that would identify an entry as a duplicate and automatically insert the value of the autoincrement key when data is committed.
| Update | Merge | Resulting Action |
|---|---|---|
| false | false | Ignore row on duplicate |
| true | false | Update only values whose corresponding columns are currently NOT NULL in the database |
| false | true | Update only values whose corresponding columns are currently NULL in the database |
| true | true | Update all values to new data |
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Initiate a SqlDataManager object.
Before adding data to the SqlDataManager, you must build the schema of any tables you will use, as well as add foreign keys if you are not using a database engine that natively supports them (such as InnoDB for MySQL).
Returns a SqlDataManager. If an error is experienced it will be thrown.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
com.screenscraper.datamanager.sql.SqlDataManager
SqlDuplicateFilters are designed to filter duplicates when more information than just a primary key might define a duplicate entry. For example, you might define a unique person by their SSN, driver's license number, or by their first name, last name, and phone number. It is also possible that a single person may have multiple phone numbers, and if any of them match then the duplicate constraint should be met. Using an SqlDuplicateFilter can check for conditions such as this and correctly recognize duplicate entries.
This feature is only available for Professional and Enterprise editions of screen-scraper.
Sometimes the data will need to be filtered across multiple tables, or possibly different constaints might indicate a duplicate. An example of this is a person might be a duplicate if their SSN matches OR if their driver's license number matches. Alternatively, they may be a duplicate when they have the same first name, last name, and phone number.
Duplicate filters are checked in the order they are added, so consider perfomance when creating duplicate filters. If, for instance, most duplicates will match on the social security number, create that filter before the others. Also make sure to add indexes into your database on those columns that you are selecting by or else performance will rapidly degrade as your database gets large.
Duplicates will be filtered by any one of the filters created. If multiple fields must all match for an entry to be a duplicate, create a single filter and add each of those fields as constraints, as shown in the third filter created above. In other words, constraints added to a single filter will be ANDed together, while seperate filters will be ORed.
Add a constraint that checks the value of new entries against the value of entries already in the database for a given column and table.
Returns void.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Sometimes the data will need to be filtered across multiple tables, or possibly different constaints might indicate a duplicate. An example of this is a person might be a duplicate if their SSN matches OR if their driver's license number matches. Alternatively, they may be a duplicate when they have the same first name, last name, and phone number.
Duplicate filters are checked in the order they are added, so consider perfomance when creating duplicate filters. If, for instance, most duplicates will match on the social security number, create that filter before the others. Also make sure to add indexes into your database on those columns that you are selecting by or else performance will rapidly degrade as your database gets large.
Duplicates will be filtered by any one of the filters created. If multiple fields must all match for an entry to be a duplicate, create a single filter and add each of those fields as constraints, as shown in the third filter created above. In other words, constraints added to a single filter will be ANDed together, while seperate filters will be ORed.
Create an SqlDuplicateFilter for a specific table and register it with the data manager.
Returns an SqlDuplicateFilter that can then be configured for duplicate entries.
| Version | Description |
|---|---|
| 5.0 | Available for professional and enterprise editions. |
Oftentimes you want to write extracted data directly to an XML file. This class facilitates doing that. Before working with the methods below, you may wish to read our documentation about writing extracted data to XML, which contains examples of scripts that utilize these methods.
This feature is only available to Enterprise editions of screen-scraper.
Initiate a XmlWriter object.
Returns a XmlWriter. If an error is experienced it will be thrown.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
| 5.5.3a | Added the constructor that takes a character set. |
com.screenscraper.xml.XmlWriter
Add a node to the XML file.
Returns the added element object.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Add multiple nodes under a single node (new or already in existence).
OR
OR
Returns the main added element object, if one was created. It there was not a main element that was added then it returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
Close the XmlWriter.
This method does not receive any parameters.
Returns void.
| Version | Description |
|---|---|
| 4.5 | Available for enterprise edition. |
The REST API was first released in the stable version 5.0 (alpha 4.5.18a). It is not a true REST API but rather an API accessible via GET requests. But for the sake of naming we call it the screen-scraper REST API. It will allow you to issue web interface commands through GET requests.
The basic structure to all REST API requests is to specify the action GET parameter with what you want to do. Some actions will require other parameters to be set as well. Here are some available actions and their parameters.
For any of this to work screen-scraper has to be running in server mode.
This feature is only available to Enterprise editions of screen-scraper.
The returned file now contains the scrapeable_session_id of the scrape to ease in manipulating it with other REST Interface actions.
All requests require that you pass your registered email address, which will be determined when you sign up for the anonymization service. This is passed as a URL-encoded string in the URL query string using the key registered_email. Your password will also be required, which is passed to the server via the password parameter.
Each call to the server is done via a GET request. The possible requests are described below:
Expect an average delay of around 20 seconds before receiving a response from the system for reach request made.
Here's an example of what would be returned from this request:
ec2-75-101-238-93.compute-1.amazonaws.com:3128 i-61955e08
ec2-75-131-250-53.compute-1.amazonaws.com:3128 i-6e955e07
Each proxy gets its own line. The host and port are given first, then a space character, then the instance ID.
You'll use the instance ID if you want to report a proxy as bad (so that it will be terminated and one will be spawned in its place).
After terminating a proxy, it will take a minute or two to spawn one in its place. You'll want to query the server periodically in order to refresh your current pool of proxies.
When writing scripts within screen-scraper, there are a number of objects and methods available to you. You can view the stable objects and classes available to scripts in the API section of our documentation. This sections only documents those methods that are in a current alpha release. You are welcome to use them but know that they are prone to change. We always work for backwards compatibility of stable features but with alpha features we will not guarantee compatibility until they appear in a stable version.
Alpha methods and objects are only available if you have screen-scraper upgrade to unstable versions. We don't guarantee that the methods will not change after their introduction if improvements are required, desired, or purposes change.
The examples are given using Interpreted Java as the scripting language. This is in accordance with the stable API.
Applies an XPath expression to the current HTML response. If tidying the response failed this method will also fail.
An XmlNode. See example for usage.
| Version | Description |
|---|---|
| 6.0.1a | Available in Professional and Enterprise editions. |
This feature is only available to Professional and Enterprise editions of screen-scraper.
screen-scraper was designed from the beginning to interact with external systems. This means that you can invoke it from other applications and it can send data to yet other systems. We've tried to design screen-scraper such that it can interact with code written in virtually any modern programming language or platform. This section of our documentation will familiarize you with how this process occurs, as well as specifics on languages that screen-scraper can work with.
In order to interact with screen-scraper in any of these methods it needs to be running in server mode.
Read, write and query your database from within screen-scraper. SqlDataManager object or traditional approach.
In order to invoke screen-scraper from something like a Visual Basic application or a PHP script screen-scraper needs to be run in server mode. When running screen-scraper in server mode it acts, in many respects, like a database server would. Interacting with screen-scraper via one of the remote scraping sessions is more or less analogous to querying a database via a database driver. Objects like scrapeable files and scripts set up in screen-scraper provide access to information much like tables and columns in a database would. Issuing a SQL query to a database is analogous to setting session variables and telling screen-scraper to initiate a named scraping session.
When running as a server any error messages screen-scraper produces will be written to the error.log file, found in the log folder of screen-scraper's install directory.
Each time you run a scraping session externally screen-scraper will generate a log file corresponding to that scraping session in the log folder found inside screen-scraper's install directory. This can be invaluable for debugging, so you'll want to take a look at it if you run into trouble.
You can turn server logging off by unchecking the Generate log files check box in the Servers section of the settings dialog box. Within a scraping session script, the name of the log file can be found in the session variable SS_LOG_FILE_NAME.
Once the server is running it will be listening for connections. The default port the server will listen on is 8778, which can be changed by altering the Port in the Servers section of the settings dialog box.
While the server is running scraping sessions and scripts can be imported into screen-scraper without stopping it. See the documentation on importing and exporting objects for more information.
If you are experiencing trouble trying to connect to screen-scraper you will want to check the connection restrictions.
To interact with screen-scraper from ColdFusion we will make use of ColdFusion's ability to implement Java classes. The ColdFusion server can execute Java code but needs to be configured first.
screen-scraper needs to be running as a server before invoking it externally.
These steps assume that the server is running on the local machine and default setup. Some consideration might need to be given if you have changed these settings.

Now that the server will be able to load the Java class and interact with the screen-scraper server through it, you now need to learn the methods and objects. As you are really interacting through Java you will need to read the section on Invoking screen-scraper from Java to learn about the methods and classes.
The following is an example of a .cfm file that creates a RemoteScrapingSession object, calls the scrape method and processes the results. The scraping session to be invoked is called test and we will be outputing a session variable, TEST that we have saved in screen-scraper.
For another example of interacting with screen-scraper via ColdFusion please see Tutorial 4: Scraping a Shopping Site from External Programs.
A Java application or servlet interacts with screen-scraper via the class RemoteScrapingSession class (com.screenscraper.scraper.RemoteScrapingSession). You can utilize the class by including the screen-scraper.jar and lib\log4j.jar files in your CLASSPATH.
screen-scraper needs to be running as a server before invoking it from a Java class.
You can also reference your own Java code from within screen-scraper
The following is a reference for all of the methods found in the RemoteScrapingSession class.
Currently only Strings, DataRecords, and DataSets can be accessed by this method.
The server cannot be started remotely.
It is also possible to store data sets and data records in session variables, which can then be accessed via the RemoteScrapingSession class. Data set objects are analogous to database result sets and data records are analogous to individual records within a result set. When an extractor pattern is applied a data set is generated. Storing the resulting data set in a session variable (within a screen-scraper script) will allow for it to be accessed via a RemoteScrapingSession.getVariable call. More information on these classes can be found in the DataRecord and DataSet API documentation pages.
This feature is only available to Enterprise editions of screen-scraper.
The DataReceiver (com.screenscraper.scraper.DataReceiver) interface allows your code to handle extracted data as it is being scraped. That is, you need not wait until the scraping session has finished before getting access to the extracted data. This interface contains a single method:
Once you have implemented the DataReceiver interface on any of your own classes, then pass an instance of the class to the RemoteScrapingSession via the setDataReceiver method. Here are other methods that allow you to control the flow of real time information.
On the screen-scraper side, whenever you'd like to send data from screen-scraper back to your code, you simply invoke the session.sendDataToClient method. Data sent through this method will show up through the receiveData method.
As a specific example, let's suppose you've created a scraping session that extracts product records from a shopping web site. As each product record is being scraped, you might simply output them to a CSV file, but you decide instead that you'd like to insert them into your database, and determine that it would be best for you to write your own code to perform the database insertion. In your scraping session, you might have a script that contains the following:
You set up this script to be invoked After each pattern match for the extractor pattern that pulls the product information. For example, the extractor pattern might get the price, title, and weight of the product. Because the script is being invoked After each pattern match, the current dataRecord object will hold all of that information. You invoke session.sendDataToClient so that each record can be processed by your code as it gets extracted.
In your Java code you create a class that implements the DataReceiver interface. You create an instance of this class and pass it to your RemoteScrapingSession object so that you can process each of the product records as they get extracted. Your >receiveData method implementation might look something like this:
Each time you invoke session.sendDataToClient in screen-scraper, there will be a corresponding method call made to your receiveData method, which will allow you to handle each of the data pieces individually.
For other examples of using the Java driver please see Tutorial 3: Extending Hello World and Tutorial 4: Scraping a Shopping Site from External Programs.
A PHP script interacts with screen-scraper via a PHP class called RemoteScrapingSession. You can utilize this class by including the file remote_scraping_session.php (found in the misc/php directory of your screen-scraper installation) within your PHP script.
screen-scraper needs to be running as a server before invoking screen-scraper from a PHP script.
The following is a reference for all of the methods found in the RemoteScrapingSession class.
Currently only Strings, DataRecords, and DataSets can be accessed by this method.
Calling this method will only have an effect if it's done before calling the scrape method. If this value is set to true, after the scrape method is called, program flow will return immediately, but the scraping session will still be running in screen-scraper.
This feature is only available to Enterprise editions of screen-scraper.
By creating a special PHP class, your code can handle extracted data as it is being scraped instead of after the scrape is finished. That means, you will not need to wait until the scraping session has finished before getting access to the extracted data.
We recommend calling the class DataReceiver.
The DataReceiver class needs to contain the following method (you can add other methods as needed to process the data but this one is particular).
Once you have created the DataReceiver class containing the receiveData method it must be incorporated into the RemoteScrapingSession using the setDataReceiver method. Here are other methods that allow you to control the flow of real time information.
On the screen-scraper side, whenever you'd like to send data from screen-scraper back to your code, you simply invoke the session.sendDataToClient method. Data sent through this method will be processed through the receiveData method.
As a specific example, let's suppose you've created a scraping session that extracts product records from a shopping web site. As each product record is being scraped, you might simply output them to a CSV file, but you decide instead that you'd like to insert them into your database, and determine that it would be best for you to write your own code to perform the database insertion. On the screen-scraper side, in your scraping session, you might have a script that contains the following:
You set up this script to be invoked After each pattern match for the extractor pattern that pulls the product information. For example, the extractor pattern might get the price, title, and weight of the product. Because the script is being invoked After each pattern match, the current dataRecord object will hold all of that information. You invoke session.sendDataToClient so that each record can be processed by your code as it gets extracted.
In your PHP code you create a class that implements the receiveData( $key, $value ). You create an instance of this class and pass it to your RemoteScrapingSession object so that you can process each of the product records as they get extracted. Your DataReceiver class implementation might look something like this:
You would instantiate the class and set it on your session like so:
Each time you invoke session.sendDataToClient in screen-scraper, there will be a corresponding method call made to your receiveData method, which will allow you to handle each of the data pieces individually.
For other examples of using the PHP driver please see Tutorial 3: Extending Hello World and Tutorial 4: Scraping an E-commerce Site from External Programs.
A Python script interacts with screen-scraper via a Python class called RemoteScrapingSession. You can utilize this class by importing the module remote_scraping_session.py (found in the misc/python directory of your screen-scraper installation) within your Python script.
screen-scraper needs to be running as a server before invoking screen-scraper from a Python script.
The following is a reference for all of the methods found in the RemoteScrapingSession class.
Currently only Strings, DataRecords, and DataSets can be accessed by this method.
Calling this method will only have an effect if it's done before calling the scrape method. If this value is set to true, after the scrape method is called, program flow will return immediately, but the scraping session will still be run by screen-scraper.
For an example of using the Python driver please see Tutorial 4: Scraping a Shopping Site from External Programs.
A Ruby script interacts with screen-scraper via a Ruby class called RemoteScrapingSession. You can utilize this class by importing the module remote_scraping_session.rb (found in the misc/ruby directory of your screen-scraper installation) within your Ruby script.
screen-scraper needs to be running as a server before invoking screen-scraper from a Ruby script.
The following is a reference for all of the methods found in the RemoteScrapingSession class.
Currently only Strings, DataRecords, and DataSets can be accessed by this method.
Calling this method will only have an effect if it's done before calling the scrape method. If this value is set to true, after the scrape method is called, program flow will return immediately, but the scraping session will still be running in screen-scraper.
For an example of using the Ruby driver please see Tutorial 4: Scraping a Shopping Site from External Programs.
When running as a server screen-scraper can be invoked from any windows application that supports COM, such as Visual Basic, Active Server Pages, or Visual C++. For examples of using the COM driver please see Tutorial 3: Extending Hello World and Tutorial 4: Scraping a Shopping Site from External Programs.
In order to use the COM driver with screen-scraper you'll need the Microsoft Virtual Machine installed on your system (not Sun's Java Runtime Environment). It's likely you've already got it on your computer, but if you experience problems you should probably try installing the most recent version of the Microsoft Virtual Machine, which can be downloaded from java or jheroen.
When screen-scraper was installed it registered the COM driver on your system in the form of a DLL.
It's very likely that you need not do anything further in order to make use of the DLL.
Should you run into trouble, though, you might try re-registering the DLL:
A Windows-based application interacts with screen-scraper via the DLL mentioned previously (think of it as a database driver).
screen-scraper needs to be running as a server before being invoked via the COM driver.
The following is a reference for all of the methods of the RemoteScrapingSession COM object.
StoreVariable method. Calling this method will cause the DLL to release from memory the data set or data record identified by VariableName.
Scraping sessions created within screen-scraper can be invoked by running screen-scraper from a Unix terminal or a DOS command prompt. This allows for possibilities such as scraping information at regular intervals via something like cron or a scheduled task. The basic syntax is as follows:
If you installed a version of screen-scraper that includes a Java Virtual Machine (currently Windows and Linux), you'll want to preface the command with "jre\bin\" on Windows or "jre/bin/" on Linux.
You could also do it in two steps. In which case the two commands are represented below.
{screen-scraper-install-folder} is the location where you installed screen-scraper, such as "C:\Program Files\screen-scraper professional edition\".
This would invoke the Google search scraping session and pass in a parameter named search_string containing the value screen scraper. This will cause a session variable named search_string to be created, which would hold the value screen scraper.
Passed-in parameters need to be URL-encoded strings, just like the query string in a URL.
This one would invoke the Hotmail mail retrieval scraping session and pass in two parameters: user_name containing the value uname and password containing the value mypass. These parameters will become session variables.
While running screen-scraper from command line, you can have the log written to a file by piping it. In order to do that, you need to change the code from the above examples.
For the first example, lets say you want to write a log file with a name google_search.log, the code would change to:
The only difference is at the end of the request: > "log\google_search.log". This instructs the log of the scrape to be written to the log\google_search.log file.
The bat file with the above code needs to be inside the folder where screen-scraper is installed. But if you want your bat file somewhere else other than the screen-scraper installed directory, you have to make some changes to the code. First, you have to cd to the directory where screen-scraper is installed. The code will look like this:
Similarly, for the second example the code to write a log file will be:
The above code will write a log file hotmail_mail_retrieval.log inside the log directory.
If your bat file is not inside the screen-scraper installed directory, the code should be like this:
When running on Windows, any % character needs to be doubled because this character is treated in a special way in DOS. For example, the parameter "string=hello%21world" would need to be passed in as "string=hello%%21world".
While running screen-scraper from command line, there is one thing we need to consider: Memory size. Java runs with a fixed amount of heap memory, which happens to be 64Mb by default. If you get an error message that says it's out of memory then this is because screen-scraper consumed all the heap memory and requires more in order to continue its job.
You can increase the heap memory with the -Xmx flag. To set the heap memory size to 1024 megabytes, use the flag below.
Lets say, we got an error message out of memory size, while running the Hotmail mail retrieval scraping session (from the examples above). The code to increase the heap memory size will be:
This code will increase the heap memory size of java to 1024 megabytes.
Remember not to set the heap memory size larger than the physical memory of the machine you are running on.
This feature is only available to Enterprise editions of screen-scraper.
SOAP is a common protocol used for accessing web services based on XML. There are several libraries available in most popular programming languages which allow for the rapid development of SOAP clients.
Many of the libraries available include some method of generating the code necessary to interact with a specific SOAP interface when given a WSDL file. We have also provided two examples using a SOAP client for screen-scraper: Java and .NET.
string getLog(string filename) - Return the content of a given log file.
string getLog(string filename, boolean start, int lines) - Returns a portion of the content of a given log file.
string[] getLogNames() - Returns the names of all the files in the log directory of the remote server.
long getLogSize(string filename) - Return the size of the given logfile in bytes.
int removeLog(string filename) - Remove a log file from the log directory on the remote server.
string[] getCompletedScrapingSessions() - Returns the ID's of the completed scraping sessions.
string[] getDataRecord(string id, string var) - Get a data record for the given variable.
string[][] getDataSet(string id, string var) - Get the data set contained in a variable in a scraping session.
string[] getRunningScrapingSessions() - Return the ID's of the currently running scraping sessions.
string getScrapingSessionName(string id) - Returns the name of the scraping session where its key is id.
string[] getScrapingSessionNames() - Returns an array of names of scraping sessions which this server currently has.
long getScrapingSessionStartTime(string id) - Returns the starting time of a particular scraping session as a long.
string[] getScriptNames() - Returns the names of scripts in this server.
string getVariable(string id, string var) - Get the value of a certain variable in a scraping session.
string initializeScrapingSession(string name) - Initialize this scraping session to allow it to be scraped.
int isFinished(string id) - Returns if the session with key=id is finished.
int removeCompletedScrapingSession(string id) - Remove the scraping session given by id from the list of completed scraping sessions.
int removeScrapingSession(string name) - Remove a scraping session from the remote server and from it's database.
int removeScript(string name) - Remove a script from the remote server and it's database.
int scrape(string id) - Scrape the session given by this ID.
int setTimeout(string id, int minutes) - Set the time out minutes of a scraping session to scrape.
int setVariable(string id, string var, string value) - Set a variable within a scraping session.
int stopScrapingSession(string id) - Stop a scraping session in progress.
int update(string xml) - Update the remote server with an exported scraping session or script.
boolean isAcceptingConnections() - Returns the value to acceptingConnections, which is the value which dictates if the server is handling remote requests to scrape.
int setAcceptingConnections(boolean accepting) - Sets the value for acceptingConnections, which will either stop the server from handling requests for remote scrapes or allow them.
public static boolean isAcceptingConnections()
Returns the value to acceptingConnections, which is the value which dictates if the server is handling remote requests to scrape.
Returns: true if the server is will accept requests to scrape.
public static int setAcceptingConnections(boolean accepting)
Sets the value for acceptingConnections, which will either stop the server from handling requests for remote scrapes or allow them.
Returns: int which represents success or a specific error code.
public string[] getScrapingSessionNames()
Returns an array of names of scraping sessions which this server currently has.
Returns: names of scraping sessions.
public string[] getScriptNames()
Returns the names of scripts in this server.
Returns: names of scripts.
public string[] getRunningScrapingSessions()
Return the ID's of the currently running scraping sessions.
Returns: An array of Strings, which are the ID's.
public string[] getCompletedScrapingSessions()
Returns the ID's of the completed scraping sessions. (Also, updates the list.)
Returns: the ID's of completed scraping sessions.
public int removeCompletedScrapingSession(string id)
Remove the scraping session given by id from the list of completed scraping sessions.
Returns: an int representing success or a failure code.
public int isFinished(string id)
Returns if the session with key=id is finished.
Returns: an int representing finished (1), not finished (0) or error (0)
public string getScrapingSessionName(string id)
Returns the name of the scraping session where its key is id.
Returns: the name of a scraping session, or "-1" if not found.
public long getScrapingSessionStartTime(string id)
Returns the starting time of a particular scraping session as a long.
Returns: the starting time of the scraping session, -1 if not yet started, or 0 if session not found.
public string initializeScrapingSession(string name)
Initialize this scraping session to allow it to be scraped.
Returns: if success then the ID of this scraping session is returned, otherwise "-1".
public int scrape(string id)
Scrape the session given by this ID.
Returns: 0 if an error occurred or 1 if successfully started.
public int setVariable(string id, string var, string value)
Set a variable within a scraping session. Disallowed if acceptingConnections is false.
Returns: 1 if successfully set, 0 otherwise.
public int setTimeout(string id, int minutes)
Set the time out minutes of a scraping session to scrape.
Returns: 1 if successful, 0 otherwise.
public int stopScrapingSession(string id)
Stop a scraping session in progress.
Returns: 1 if successful, 0 otherwise.
public string getVariable(string id, string var)
Get the value of a certain variable in a scraping session. Note that currently only Strings, DataRecords, and DataSets can be accessed by this method.
Returns: if this is a valid scraping session and the value of this variable is a string, then - the value is returned, "NULL" if the value is null, and "-1" otherwise.
public string[] getDataRecord(string id, string var)
Get a data record for the given variable.
Returns: an array of String objectss like key=value or an empty array if an error happened or the variable is empty.
public string[][] getDataSet(string id, string var)
Get the data set contained in a variable in a scraping session.
Returns: an array of data records as translated to arrays of String objects.
public int update(string xml)
Update the remote server with an exported scraping session or script. As a warning, if the version of screen-scraper this xml was exported from is different from the version of screen-scraper which is running as a server, then the update may not work.
Returns: 0 for failure, 1 for success.
public int removeScrapingSession(string name)
Remove a scraping session from the remote server and from it's database.
Returns: 0 for failure, 1 for success.
public int removeScript(string name)
Remove a script from the remote server and it's database.
Returns: 0 for failure, 1 for success.
public string[] getLogNames()
Returns the names of all the files in the log directory of the remote server.
Returns: an array of the names of the log files, or null - if there is no log directory.
public long getLogSize(string filename)
Return the size of the given logfile in bytes.
Returns: a long representing the length in bytes of this file, or 0 if the file - does not exist or is empty.
public string getLog(string filename)
Return the content of a given log file.
Returns: a String of the contents of the file, or "" if not possible.
public string getLog(string filename, boolean start, int lines)
Returns a portion of the content of a given log file.
Returns: a portion of the content of the given log file, or "" if anything goes wrong.
The .NET SDK includes an executable that can automatically generate the files necessary to access screen-scraper's SOAP interface as an object.
The first step in the process is getting screen-scraper running as a server.
Next we will generate the service class to do the actual communication in SOAP for us. There is a wsdl.exe in the Bin directory of the .NET SDK. Find it on your computer. Using v1.1 type this command:
To see the options available when using wsdl.exe like the output language being Visual Basic, try the flag /?.
After creating the SOAPInterfaceService class, it is possible that there is a mistake in the code. Find the getDataSet method. If the method returns String[], then change it to String[][] and also the casting of the returned object.
The following is an example class which uses the generated class from above to call on the scraping session created in Tutorial 2.
Be sure that the newly created class is part of the compilation process.
The Axis Library included with screen-scraper makes creating a SOAP client quite easy. Using the WSDL file created from the remote procedures on the screen-scraper server, Axis can create the stubs to make calling the methods in SOAP a matter of just using an object.
The first step in the process is getting screen-scraper running as a server.
The next step is to call the WSDL2Java class in the Axis library. To do so there are six jar files which need to be in the class path to run the class. All of these files are included in the lib directory of screen-scraper.
You also need to call the correct class org.apache.axis.wsdl.WSDL2Java. We also recommend changing the output package that the created Java files go to so that it is com.screenscraper.soapclient. This can be done using the --package option.
Here is an example command-line usage of options just specified above from inside the screen-scraper directory on Windows.
This will create a new directory com where the command was issued containing the necessary stubs.
The following is an example class which uses the generated classes from above to call on the scraping session created in Tutorial 2.
Be sure that the newly created files compile with this one and that the above mentioned jars are in your CLASSPATH. Also, make sure that screen-scraper is running as a server.
If you're using Visual Studio 2008 or later, the project 'Target Framework' will need to be set to .NET 3.5 or later. However, do not use any .NET client frameworks since they do not have the required libraries for your project to compile.
A C# application interacts with screen-scraper via the Screenscraper.RemoteScrapingSession class. You can utilize the this class by compiling with a reference to the misc/dotNET folder of your screen-scraper distribution.
screen-scraper needs to be running as a server before invoking it from a .NET class.
The following is a reference for all of the methods found in the RemoteScrapingSession class.
It is also possible to store data sets and data records in session variables, which can then be accessed via the RemoteScrapingSession class. Data set objects are analogous to database result sets and data records are analogous to individual records within a result set. When an extractor pattern is applied a data set of data record objects is generated. Storing the resulting data set in a session variable (within a screen-scraper script) will allow for it to be accessed via a RemoteScrapingSession.GetVariable call.
The data record class (Screenscraper.DataRecord) simply extends Microsoft's Hashtable.
The following is a reference for all of the methods found in the DataSet class (Screenscraper.DataSet).
For an example of using the .NET driver please see Tutorial 4: Scraping an E-commerce Site from External Programs.
This section s provided to give additional information about the software, how it works, and the technologies behind screen-scraper. Many of these pages contain links that are not under the control of our company. We have chosen them for their quality at the time. If the links break or the content changes we would appreciate your contacting us about it so that the links remain relevant.
As of screen-scraper 5.0 a simple code completion has been added to the scripting. It is meant to make it easier to remember the method names and their parameters. It provides you with this information as well as a link back to the documentation on the methods.
To activate the dialog simply type the name of a built-in object followed by a period (just like you would when coding). If you pause after the period the dialog will pop up and allow you to click through the methods of the object. As you type it will limit the list until it gets to the one that you are looking for. By double-clicking on one, or hitting the tab key when it is selected, you will get the remaining code in your script with place holders for the parameters. Type in the values of the parameters and hit tab to jump to the next. When you are finished, the last time you hit tab it will jump to the end of the method call.
In addition to the code completion there are a number of built-in macros for common tasks. To active a macro simply type in its code and then hit the spacebar while holding down the Ctrl button.
Overview
This feature is only available to the Enterprise edition of screen-scraper.
screen-scraper has the ability to automatically generate RSS and Atom feeds from extracted data. If you're unfamiliar with RSS and Atom feeds you might take a minute to read up on the topic first.
The documentation on this page is a bit abstract. If you're interested in building RSS/Atom feeds with screen-scraper it would probably be a good idea for you to go through our Sixth Tutorial, which will walk you through the process in detail.
How it Works
A small web server runs within screen-scraper that interacts with the scraping engine. As such, you can access a URL within a browser or RSS/Atom reader that will cause screen-scraper to invoke a scraping session, then return back an RSS or Atom feed.
The basic syntax for the URL you'll use to generate a feed looks like this:
For example, if you were running screen-scraper on your local machine, and wanted to generate a feed for the "Shopping Site" example used in our tutorials with the search term "bug" the URL would look like this:
As with any other URL, each of the parameters must be properly URL-encoded. Key/value pairs can also be passed in as POST parameters.
The only required parameter is "scraping_session". screen-scraper will create session variables out of any other parameters that get passed in.
Setting Up the Scraping Session
The scraping session must have certain named elements present in order to generate the feed. They are as follows:
When the XML feed is requested through your browser or reader screen-scraper will invoke the scraping session named by the "scraping_session" parameter. Once the scraping session completes screen-scraper will look for a DataSet called "XML_FEED", iterate over its constituent DataRecord objects, building the feed from them.
Hypertext transfer protocol provides a way for clients such as web browsers to communicate with web servers. There's quite a bit on the web that's written on the topic, so for the time being we'll just provide some good links for you:
Scraping sessions and scripts can be exported from screen-scraper to external files. You might consider doing this in order to back up your work, and even commit them to a versioning system, such as CVS or Subversion.
In order to export a scraping session or script to an external file simply select the object you wish to export then click on the corresponding Export button (Export Session or Export Script). You'll be asked to save the file to a location of your choice. You're also free to name the file what you wish, though we recommend you leave the (scraping session) or (script) portion of the name in tact so that you can identify the type of the object later on. When you export a scraping session from screen-scraper all scripts directly associated with that scraping session will be exported within the same file.
When a scraping session is exported the time of export is also included in the resulting file. This date can be useful to track versions of the scraping session. To view the date, open the .sss file in a text editor and search for the
To import a scraping session or script into screen-scraper select the Import... option from the . Locate the ".sss" file corresponding to the object you wish to import, and select Open. If you've selected a valid file the objects contained within that file will be imported into the application.
You can also import exported scraping sessions and scripts into screen-scraper by copying them into the import folder you'll find in the directory where screen-scraper was installed. This can be especially useful while screen-scraper is running as a server, which allows the objects to be imported on the fly (that is, without stopping the server). screen-scraper will check this directory just before executing a scraping session, and import any files found in it. Note that imported files will be removed from the import folder once they are imported by screen-scraper.
In cases where you want to pack up scraping sessions and scripts along with other files needed to run a scrape, you can compress them all into an update.zip file. This file should replicate the directory structure of screen-scraper. For example, you might have a folder called import that contains a scraping session. You might also have a CSV file in the root of the zip file that contains parameters needed to run the scraping session. You can zip all of these up into an update.zip file, then place that file inside an update folder found in screen-scraper's install directory. When screen-scraper starts up it will unzip the file, copy all of its contents to the corresponding locations, then delete the update.zip file.
If you've un-checked the Overwrite on import checkbox for a script, and would like to import that script into an instance of screen-scraper that is running in a GUI-less environment, follow the instructions on script overwriting.
The memory usage indicator was introduced in screen-scraper 4.5 and shows you how much of the memory currently allocated to screen-scraper that is being used. As screen-scraper requires it, it may be allocated more memory from the underlying Java Virtual Machine, up to the amount specified in the settings dialog box.
In the workbench, the indicator is on the far right of the main window's status bar. In the Enterprise Edition's web interface, the indicator is at the top of the page under the Import button.
The current memory usage can also be queried in a script via the getMemoryUsage method.
Often times screen-scraper will be running on a server that has no graphical interface. Updating to the latest version in such an environment previously required multiple steps, but can now be done with a simple Python script.
You can download the script from our site.
Any Unix-based computer worth its salt will already have python installed. To use the updater, open a terminal and navigate to the screen-scraper install directory. Ensure that screen-scraper is not currently running (via ./server status). After that, issue this command to update to the latest version:
If you want to force screen-scraper to upgrade to the latest unstable version, use this command:
Alpha versions are used to fix minor bugs and feature enhancement testing before they are added to stable versions. As such anything that is in the alpha version is prone to change and instability as they are being improved. Because alpha versions are not considered production-quality stable, we recommend you back up your work before upgrading. That said, we're grateful for any willing to venture into alpha territory to help us test. Please let us know when you find bugs and such.
See also: Alpha API.
Alpha versions are used to fix minor bugs and feature enhancement testing before they are added to stable versions. As such anything that is in the alpha version is prone to change and instability as they are being improved. This log will follow the changes as they are made for your convenience.
View Release Notes for public versions.