Download latest jsoup jar file (Download Link).
Compile code with appropriate class path value, like
javac -cp "C:\jsoup-1.7.1.jar" "TestClass.java"
java -cp "C:\jsoup-1.7.1.jar" TestClass
Simple Example using Jsoup to connect to server using login credentials and then retrieving specific page.
[sourcecode language="java"]
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.Connection;
import java.io.IOException;
import org.jsoup.Connection.Method;
import java.util.HashMap;
import java.util.Map;
import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.select.Elements;
public class TestClass
{
public static void main(String args[]) throws IOException
{
Document doc = Jsoup.connect("<URL>").get();
Elements viewState = doc.select("input[name=__VIEWSTATE");
Elements eventValidation = doc.select("input[name=__EVENTVALIDATION]");
Map<String,String> allFields = new HashMap<String,String>();
allFields.put("__VIEWSTATE", viewState.val());
allFields.put("__EVENTVALIDATION", eventValidation.val());
allFields.put("txtLogin", "<USERNAME>");
allFields.put("txtPassword", "<PASSWORD>");
allFields.put("butSubmit", "Sign In");
System.out.println(allFields);
Connection.Response res = Jsoup.connect("<URL2>")
.userAgent("Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/535.21 (KHTML, like Gecko) Chrome/19.0.1042.0 Safari/535.21")
.data(allFields)
.method(Method.POST).
execute();
String sessionId = res.cookie("<COOKIENAME>");
System.out.println(sessionId);
Document doc2 = Jsoup.connect("URL3")
.cookie("ASP.NET_SessionId", sessionId)
.timeout(0)
.get();
System.out.println(doc2.html());
}
}
[/sourcecode]
Showing posts with label Text Mining. Show all posts
Showing posts with label Text Mining. Show all posts
Wednesday, January 30, 2013
Web scrapping using Jsoup
Labels:
java,
java library,
jsoup,
Text Mining,
web scrapping
Thursday, October 20, 2011
Sentiment analysis using Naive Bayes Algorithm
Experimented with simple Naive Bayes for sentiment classification.
Naive Bayes code is available here chatper6/docclass.py and training data is available here
Changed the getwords() function in docclass.py
- to remove special characters like single-quote, comma, full stop from text
- to split based on white spaces instead of non word character because it ignored emots with non word character split and
- included nltk stopwords corpus check.
[sourcecode language="python"]
def getwords(doc):
doc=re.sub('\.+|,+|!+|\'','',doc)
splitter=re.compile('\\s+')
#print doc
# Split the words by non-alpha characters
words=[s.lower().strip() for s in splitter.split(doc)
if s.lower().strip() not in nltk.corpus.stopwords.words('english') ]
print words
# Return the unique set of words only
return dict([(w,1) for w in words])
[/sourcecode]
For training data, converted ';;' separated data file to '\t' separated file because csv.reader() function
was not accepting two symbol delimiters.
Changed sampletrain function to train classifier on training data file "testdata.manual.2009.05.25".
[sourcecode language="python"]
def sampletrain(cl):
read = csv.reader(open('pos 1', 'rb'), delimiter='\t')
cnt = 1
for row in read:
if row[0] == 0:
sent = 'bad'
else:
sent = 'pos'
data = row[5]
cl.train(data,sent)
cnt = cnt+1
print cnt
[/sourcecode]
Naive Bayes code is available here chatper6/docclass.py and training data is available here
Changed the getwords() function in docclass.py
- to remove special characters like single-quote, comma, full stop from text
- to split based on white spaces instead of non word character because it ignored emots with non word character split and
- included nltk stopwords corpus check.
[sourcecode language="python"]
def getwords(doc):
doc=re.sub('\.+|,+|!+|\'','',doc)
splitter=re.compile('\\s+')
#print doc
# Split the words by non-alpha characters
words=[s.lower().strip() for s in splitter.split(doc)
if s.lower().strip() not in nltk.corpus.stopwords.words('english') ]
print words
# Return the unique set of words only
return dict([(w,1) for w in words])
[/sourcecode]
For training data, converted ';;' separated data file to '\t' separated file because csv.reader() function
was not accepting two symbol delimiters.
Changed sampletrain function to train classifier on training data file "testdata.manual.2009.05.25".
[sourcecode language="python"]
def sampletrain(cl):
read = csv.reader(open('pos 1', 'rb'), delimiter='\t')
cnt = 1
for row in read:
if row[0] == 0:
sent = 'bad'
else:
sent = 'pos'
data = row[5]
cl.train(data,sent)
cnt = cnt+1
print cnt
[/sourcecode]
Labels:
naive bayes,
python,
sentiment analysis,
Text Mining
Subscribe to:
Posts (Atom)