Showing posts with label python. Show all posts
Showing posts with label python. Show all posts
Wednesday, May 27, 2009
Python module of the week
Someone has been doing a tutorial of a python module per week for quite some time
Wednesday, March 25, 2009
Read multiple newline types in csv
One really convenient module I use all the time is the csv module. You would think that parsing csv by yourself is easy, i.e., if you've never tried it, but it's really rather tricky.
People who use excel (sneer) often need to send me csv-formatted spreadsheets, and quite often the formatting is quite messed up. Often the line-endings are bizarre and frightening, and unrecognizable.
The solution is to add a 'U' flag to the open() call to enable universal newline support.
For instance, if you're opening a file for reading:
open('somefile.csv', 'rU')
will not care what the newline format is. Somehow it manages it all without me having to think about it. My thoughts are expensive.
People who use excel (sneer) often need to send me csv-formatted spreadsheets, and quite often the formatting is quite messed up. Often the line-endings are bizarre and frightening, and unrecognizable.
The solution is to add a 'U' flag to the open() call to enable universal newline support.
For instance, if you're opening a file for reading:
open('somefile.csv', 'rU')
will not care what the newline format is. Somehow it manages it all without me having to think about it. My thoughts are expensive.
Tuesday, March 24, 2009
Monday, February 16, 2009
Disposable variables
Ever have a bunch of variables you get returned in a tuple, and you only want some of them? So you assign them to some useless variable like 'nothing' or 'deleteme' and hope that the optimizer will instantly garbage collect them away?
(user, user_type, ingorethis) = somefunction()
The correct way to do this is to use the _ variable. (underscore). Underscore is equivalent to perl's $_ variable: it contains the results of the last evaluated expression. This is sometimes useful if you're using the python shell as a calculator. So, it will instantly be overwritten.
(user, user_type, _) = somefunction()
(user, user_type, ingorethis) = somefunction()
The correct way to do this is to use the _ variable. (underscore). Underscore is equivalent to perl's $_ variable: it contains the results of the last evaluated expression. This is sometimes useful if you're using the python shell as a calculator. So, it will instantly be overwritten.
(user, user_type, _) = somefunction()
Saturday, December 20, 2008
make python understand ~ for home directories
Python does not automatically expand ~ for user home directories. There is a function in os.path that does that for you:
mypath = os.path.expanduser('~/somepath')
mypath == '/home/nick/somepath'
any and all
Python has two builtin functions for evaluating iterables for truth/falseness: any() and all(). any() returns true if any member of the iterable evaluates to True, while all returns True only if every single member of the iterable evaluates to True.
>>> any([True, False, False])
True
>>> any([False, 0])
False
>>> all([True, 0])
False
>>> all([1])
True
>>> all([1, 1, True])
True
True
>>> any([False, 0])
False
>>> all([True, 0])
False
>>> all([1])
True
>>> all([1, 1, True])
True
There are identically named subquery operators in sql, however they are only really useful for correlated subqueries which are poorly optimized and generally to be avoided in the current versions of mysql.
Friday, December 12, 2008
Parse any date easily
Use dateutil.parser.parse, possibly with the fuzzy flag set to True.
>>> from dateutil.parser import parse
>>> parse("Jan. 1st, 1991", fuzzy=True)
datetime.datetime(1991, 1, 1, 0, 0)
This is a huge boon over the tedious and underfeatured time.strptime and similar. Unhelpfully, the documentation is non-existent.
Thursday, December 11, 2008
See how python compiles some code
use the dis module to get the bytecode of any python callable.
Example:
import dis
dis.dis(lambda: x is True)
1 0 LOAD_GLOBAL 0 (var)
3 LOAD_GLOBAL 1 (True)
6 COMPARE_OP 8 (is)
9 RETURN_VALUE
Example:
import dis
dis.dis(lambda: x is True)
1 0 LOAD_GLOBAL 0 (var)
3 LOAD_GLOBAL 1 (True)
6 COMPARE_OP 8 (is)
9 RETURN_VALUE
This is another way to get the efficiency of different parts of code
Sunday, December 7, 2008
See how python encodes a string internally
Fun fact: you can see how python encodes a string internally by asking it to encode a unicode object to the 'unicode_internal' codec.
Example:
>>> u'\N{SNOWMAN}'.encode('utf16')
'\xfe\xff&\x03'
Example:
>>> u'\N{SNOWMAN}'.encode('utf16')
'\xfe\xff&\x03'
>>> u'\N{SNOWMAN}'.encode('unicode_internal')
'&\x03'
As you can see, it's UTF16 without a BOM.
Fun stuff with unicode in python
Fun fact: you can interpolate a unicode character into a unicode string literal by using \N{name of character}
Example:
>>> print u'this sign is \N{VULGAR FRACTION ONE HALF} off'.encode('utf8')
this sign is ½ off
Example:
>>> print u'this sign is \N{VULGAR FRACTION ONE HALF} off'.encode('utf8')
this sign is ½ off
You can tell the name of any given character by using the unicodedata module
>>> import unicodedata
>>> unicodedata.name(u'\xbd')
'VULGAR FRACTION ONE HALF'
The name of the famous snowman character is simply enough 'SNOWMAN.
>>> unicodedata.name(unichr(9731))
'SNOWMAN'
So you can represent this character in python simply by u'\N{SNOWMAN}'
>>> import unicodedata
>>> unicodedata.name(u'\xbd')
'VULGAR FRACTION ONE HALF'
The name of the famous snowman character is simply enough 'SNOWMAN.
>>> unicodedata.name(unichr(9731))
'SNOWMAN'
So you can represent this character in python simply by u'\N{SNOWMAN}'
Tuesday, December 2, 2008
Keep a series of python objects in a shelf
You can store a dict of pickled objects in a "shelf", a dbm-based hash table stored to disk, using the shelve module.
Sunday, November 30, 2008
Web-based auto-generated self-documentation
Fire up pydoc -p 8080, then browse to http://localhost:8080
It will generate a browseable directory of documentation for every module you have installed that's visible on sys.path .
Though the layout is hideous.
I'm not aware of another way to see every module you have installed.
Friday, November 28, 2008
Never use a [] or {} as a default argument.
I've done this a couple times.
When you make [] a default argument, it will be the same list on every subsequent invocation, with the same id(). If you append to it, what you've appended will be still be there for later invocations of the function.
Example:
def functionx(mylist = []):
When you make [] a default argument, it will be the same list on every subsequent invocation, with the same id(). If you append to it, what you've appended will be still be there for later invocations of the function.
Example:
def functionx(mylist = []):
mylist.append("F")
print mylist
for i in range(4): functionx()
The output of this program should be:
F
F, F
F, F, F
F, F, F, F
Needless to say, this is really, really stupid. The reason is because the default arguments get defined at definition time, and any mutable type will persist. This is also why somedict.setdefault({},...) will also mess up, any why defaultdict instead uses a callable.
So, you should only use constant values as default parameters, and never a mutable type. I've made this mistake a couple times.
I've seen some obtuse lunkheads actually saying it's better this way. It seems to be a rule that all language designers eventually start rationalizing their mistakes and the corners they cut (another example: see this interview with HÃ¥kon Lie , as he is unable to admit there could be any flaw in the design of CSS) and pointlessly defend bad decisions in the work they've invested so much time that no one who approached the language with fresh eyes would ever, ever defend.
I've also seen some people saying that you can use this to simulate static variables. To which I say: so, give us proper static variables then, and don't plant your language full of landmines sure to screw people up when they're not expecting it.
Wednesday, November 19, 2008
More complete regex support coming in 2.7?
Looking through the python issue tracker, looks like they're working on almost my exact wishlist of missing features in python's re module.
Including:
Including:
- possessive quantifiers, atomic grouping
- (?flag:...), (?-flag:...) support, (?flag) scoped correctly
- variable length lookbehinds
Monday, November 17, 2008
Get subpattern matches balanced even if they're optional
This will give you a mysterious "unbalanced grouping" error:
re.sub('.(.)?)', r'aaa\1')
Because the parenthetical grouping is optional, trying to access it via the \1 backreference will result in an error.
The solution is to replace the ? operator with a null alternative in the paren:
re.sub('.(|.)', r'aaa\1')
Which gives the grouping the option of containing nothing, thereby maintaining the existence of a backreference (though the backreference will of course contain an empty string, which is what you would actually expect).
re.sub('.(.)?)', r'aaa\1')
Because the parenthetical grouping is optional, trying to access it via the \1 backreference will result in an error.
The solution is to replace the ? operator with a null alternative in the paren:
re.sub('.(|.)', r'aaa\1')
Which gives the grouping the option of containing nothing, thereby maintaining the existence of a backreference (though the backreference will of course contain an empty string, which is what you would actually expect).
Wednesday, November 12, 2008
Nested list comprehensions
Fun fact: you can nest list comprehensions.
Which is exactly equivalent to:
This syntax looks like:
[(i, j) for i in range(4) for j in range(i)]
Which is exactly equivalent to:
x = []
for i in range(4):
for j in range(i):
x.append((i, j))
Monday, November 10, 2008
The hidden debug flag in the re module
The compile function of the re module, as I'm sure you already know, takes a second parameter of flags, such as re.I or re.X.
>>> re.compile(u'ell2\u1f4c', re.DEBUG)
There is an undocumented flag, however, that will print out the parse tree of the regex. This is the re.DEBUG flag.
Here's some trivial output:
>>> re.compile(u'ell2\u1f4c', re.DEBUG)
literal 101
literal 108
literal 108
literal 50
literal 8012
The 'literal' statements are followed by the ascii dec values of the characters in the pattern. I threw in a unicode character and you can see the unicode code point as the last item.
This will also tell you whether the pattern is being compiled for the first time or if it is being pulled out of the re module's cache.
Thursday, November 6, 2008
Make a more efficient list of integers
If you have a list that's all of the same type, and that type is numeric, use the array module for more efficiency.
Tuesday, November 4, 2008
Subscribe to:
Posts (Atom)