Showing posts with label python. Show all posts
Showing posts with label python. Show all posts

Thursday, December 29, 2022

pandas calculate return by shifting timeseries index

 usually, we calculate the return with next row data by simply using df.Close.shift()

but if we want shifting by a specific time period, can try:

df['ret15min'] = (df.Close.shift(-15, freq="min") -df.Close)*100/df.Close

Sunday, November 1, 2020

Install without root access / for current user

 You can use conda to install mysql client!

https://stackoverflow.com/questions/36651091/how-to-install-packages-in-linux-centos-without-root-user-with-automatic-depen/52561058#52561058

Thursday, October 29, 2020

recreate / cloning conda env with limited internet access

  • https://www.anaconda.com/blog/moving-conda-environments
  • https://docs.conda.io/projects/conda/en/latest/user-guide/tasks/manage-environments.html

Thursday, April 30, 2020

send email html content with inline image


  • preferred way:
    • script to generate html content with image encoded as base64 text
      • <img src='data:imag/png;base64,....'/>
      • base64.b64encode().decode("ascii") to get that ... string from image
    • single html file to include all the html and images, etc. 😀
    • cat that html and pipe to mutt -e "set content_type=text/html" to send the email

Monday, December 17, 2018

Where is the source file of an imported module?

from PackageABC.ParentModule import ChildModule

import inspect
print(inspect.getsourcefile(ChildModule))

#which would be in one of these paths
import sys
for p in sys.path:
    print(p)

Thursday, July 28, 2016

one-line if then else

C++
<condition> ? <if true return> : <else return>

Python
<if true return> if <condition> else <else return>

example:
rsl = f[len(prefix):] if f.startswith(prefix) else f

https://mail.python.org/pipermail/python-list/2010-July/581933.html

Sunday, January 31, 2016

How to install additional kernel to ipython / jupyter?

Assuming you already have python 3 and ipython installed with anaconda,

Create python 2.7 env
$ conda create -n py27 python=2.7

Activate the env
$ source activate py27

Install ipykernel
$ conda install notebook ipykernel

Write kernel spec
$ ipython kernelspec install-self --user
which should write to ~/.local/share/jupyter/kernels/python2

Deactivate 2.7 env
$ source deactivate

link the kernel spec to the location where ipython will read
$ cd ~/.ipython/
$ ln -s ~/.local/share/jupyter/kernels/

Confirm now that we have additional kernel
$ ipython kernelspec list

Start ipython notebook and you should now have both kernels
$ ipython notebook --no-browser

Tuesday, November 17, 2015

python pandas groupby agg percentiles?

use the ways as described in stackoverflow from google search 

or 

simply use describe:

df.groupby([col1, col2]).describe(percentiles=[.75, .95])


Optionally, you may wanna have the aggregated value display horizontally and round the numbers by appending:

df.groupby([col1, col2]).describe(percentiles=[.75, .95]).unstack().apply(lambda x:np.round(x,0))

Thursday, October 1, 2015

Friday, September 18, 2015

How to read command line output directly into pandas dataframe?

cmd = r"zgrep abc application.log | perl -pe 's/pattern/subs/'"
# python 2
pd.read_csv(StringIO.StringIO(subprocess.check_output(cmd, shell=True)))
# python 3
pd.read_csv(BytesIO(subprocess.check_output(cmd, shell=True)))

Wednesday, August 26, 2015

From MySQL to pandas df with Python 3

Install mysql connector 


# http://conda.pydata.org/docs/faq.html#id1
conda install -n <your python 3 env> mysql-connector-python

Access MySQL from python 3 with mysql connector and put result into pd df

# http://dev.mysql.com/doc/connector-python/en/connector-python-tutorial-cursorbuffered.html

import mysql.connector

# Connect with the MySQL Server
cnx = mysql.connector.connect(user='scott', database='employees')

# note that we'll have to set dictionary=True to get column name into pd and fetchall afterwards
cur = cnx.cursor(buffered=True, dictionary=True)
cur.execute('SELECT now() from dual')
pd.DataFrame(cur.fetchall())

Sunday, August 16, 2015

ipython / jupyter - how to switch kernel?

With ipython and python 2.7 installed using anaconda, how do I switch kernel to use 3.*?

$ conda create -n py34 python=3.4 anaconda
$ source activate py34
$ ipython kernelspec install-self --user
$ ipython notebook --profile=nbserver --script

Sunday, March 1, 2015

top n per group with python pandas

In [11]: df = pd.DataFrame({'cat':pd.Categorical(['A', 'A', 'A', 'B', 'B', 'B', 'C', 'C', 'C']),
   ....:                    'v1':np.random.randint(100, size=(9)),
   ....:                    'v2':np.random.randint(10, size=(9)) })

In [12]: df
Out[12]:
  cat  v1  v2
0   A  79   7
1   A  97   5
2   A  81   9
3   B  75   3
4   B  43   7
5   B  27   8
6   C  47   6
7   C  23   9
8   C  53   0

In [13]:

In [13]: # top n per group

In [14]: # 1) sort by the "top n" column (s)

In [15]: df.sort(['v1', 'v2'])
Out[15]:
  cat  v1  v2
7   C  23   9
5   B  27   8
4   B  43   7
6   C  47   6
8   C  53   0
3   B  75   3
0   A  79   7
2   A  81   9
1   A  97   5

In [16]: # 2) group by column of your choice

In [17]: # 3) select the top n of it

In [18]: df.sort(['v1', 'v2']).groupby('cat').head(2)
Out[18]:
  cat  v1  v2
7   C  23   9
5   B  27   8
4   B  43   7
6   C  47   6
0   A  79   7
2   A  81   9

In [19]: # 4) optionally sort the output nicely

In [20]: df.sort(['v1', 'v2']).groupby('cat').head(2).sort(['cat', 'v1', 'v2'])
Out[20]:
  cat  v1  v2
0   A  79   7
2   A  81   9
5   B  27   8
4   B  43   7
7   C  23   9
6   C  47   6

In [21]: