forked from yang/notes
-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathHackery.page
More file actions
3280 lines (2988 loc) · 136 KB
/
Copy pathHackery.page
File metadata and controls
3280 lines (2988 loc) · 136 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
- chef, puppet
- http://news.ycombinator.com/item?id=3090800
- android internals
- apps: bundles of components
- each app in a linux process; can be killed any time
- intents: msgs, like unix signals
- Binder: RPC system that lets services talk to components & to ea other
- IDL
- bionic: libc replacement
- dalvik vm, harmony class lib
- dx converts .class to .dex
- init: parses /init.rc; usu mounts FSs, starts mem mgr, starts services (incl
servicemanager, which manages Binder ctxs), starts zygote
- zygote aka app_process: root java process; starts system service
- system service: process that houses everything important (battery, lights,
vibrator, audio, sensors, etc), alarms, notifs, activities, windows
- toolbox: suckier busybox; freq install busybox on androids
- startup
- bootloader implements fastboot protocol; contains secure boot/recovery
logic
- no FHS but no conflicts either
- ext4 for SMP and new storage devices that were more like SD cards than NAND
flash; was yaffs2
- android development
- paths
- Environment.getExternalStorageDirectory: /mnt/sdcard
- shared space, very user-accessible, public
- generally avoid using this
- Environment.getExternalStoragePublicDirectory("AnyStringExceptNull"): /mnt/sdcard/AnyStringExceptNull
- shared space, very user-accessible, public
- may not exist; create with File.mkdirs()
- Context.getExternalFilesDir("blah"): /mnt/sdcard/Android/data/com.android.test/files/blah
- app-specific, public
- auto-creates directory
- null gives the base dir, not a subdir
- Context.getFilesDir: /data/data/com.android.test/files
- app-specific, private
- mahout
- ? means unknown
- mahout accepts csv-like files where ',' and ' ' are delims (it also ignores
the @ lines so this way it also accepts ARFF)
- linux security tools/practices
- denyhosts/fail2ban
- psad: detect intrusions from iptables
- tiger: general system audit tool
- logwatch
- chkrootkit
- ufw
- sshd_config: `PermitRootLogin no` and `PasswordAuthentication no`
- <http://www.andrewault.net/2010/05/17/securing-an-ubuntu-server/>
- machine learning libraries
- orange > pyml, mlpy
- rapidminer > weka
- weka
- trees
- REPTree: does reduced-error pruning; considers all attrs at each node
- RandomTree: no pruning; considers only random subset of attrs at each node
- J48: C4.5 impl
- <http://mydatamining.wordpress.com/2008/04/14/decision-trees/>
- dealing with dataset imbalance
- SpreadSubsample: undersampling
- SMOTE: oversampling
- CostSensitiveClassifier
- InterquartileRange: undersample (remove) outliers
- CostMatrix, Resample
- <http://old.nabble.com/Unbalanced-class-problem-td29911621.html>
- google apps script
- run on google servers
- can customize buttons/menus
- per-document, or can be added domain-wide
- can trigger on events: open, edit, install, spreadsheet form submit, etc
- can publish into public gallery or as a private or public URL service
(where it still runs *as you*)
- various actions: send mail, fetch URLs, manage docs, jdbc, utils
- speech synthesis/TTS systems
- festival: c++/scheme; basic (diphone) voices poor; HTS/Unit Selection
voices great
- flite: c
- emacspeak: dectalk 3, dectalk express, viavoice
- espeak: c
- freetts: java
- festvox: voices
- mary: sounds pretty nice, at least the online demo w cmu voices
- epos
- hts (hmm text to speech)
- mbrola: not that great
- neospeech: commercial; amazing; used in qwiki
- svox: used in google products
- process monitoring/management
- supervisord
- directly monitors children; no sub-children, no forking/start cmds
- ugly configparser files
- less control than god
- only very basic monitoring controls + event-listening extension system
- online reloading (`reread` then `update`) isn't that great
- misses certain changes, e.g. to events for event listener
- changed processes are restarted on `update`
- only one group per program
- god
- reports of problems running and memory leaks
- <http://stackoverflow.com/questions/768184/god-vs-monit>
- <http://blog.bradgessler.com/use-monit-with-rails-not-god>
- I also had problems with god hanging when playing with it
- unmaintained since Feb 2009
- monit
- nice but new language
- most actively maintained
- unlike god/supervisord:
- only supports running in fresh environment
- no stdout/stderr redirection
- only uses forking/start cmds; no child supervision
- advanced monitoring checks
- dependencies, but only process-based, not listening port-based
- janky; single-threaded architecture gets stuck waiting for pidfile to
appear and can't respond to control
- very slow/sluggish feeling
- can monitor process trees, unlike the others
- upstart
- designed for child supervision
- root user only; only supports running as root
- simple, restricted language
- no stdout/stderr redirection
- also actively maintained; ubuntu's "success story"
- dependencies
- details
- "stop on stopped foo" will stop this job even if foo just dies (doesn't
have to be explicitly stopped)
- stopping: first sends TERM, then KILL after 5s (`kill timeout 5`)
- log if exited with non-0, logged
- upstart takes care of `start on runlevel [2345]`
- don't create /etc/rc?.d/ symlinks
- ubuntu has symlinks in /etc/init.d/ to upstart-job only for compat
- perhaps users or other programs expect these there
- none of them have symlinks in /etc/rc?.d/
- nagios
- pros/features/highlights
- event processing system; events can be restarts, alerts, logging, etc.
- large ecosystem; biggest of all monitoring systems; many existing plugins
- web interface to see current status
- passive as well as active plugins; not just cron
- simple plugin API: any command line tool
- results caching, service dependencies, adaptive monitoring
- on-call notification rotations, escalations
- cons
- very easy to re-implement these checks
- yet another system to learn, with its own config labyrinth
(self-advertised as complex)
- everything is periodic, has delay
- no integration with trend monitoring systems; no data accumulation
- how is it for EC2-like places where instances are expected to fail?
- polling duration can exceed period, causing problems
- details
- define hosts, cmds, services (what cmds to run on what hosts)
- external command file: where other programs send nagios internal cmds
(not cmd-line cmds) to run
- active plugins: exit status 0-3 determine state
- distributed montoring: ssh/NRPE for active, NSCA for passive, SNMP
- passive plugins: checks _external command file_ every minute by default
- can check for sufficient freshness
- both queue results; checked every 5s by default by _reaper_
- volatile: state-triggered (v edge-triggered)
- ndoutils: store in mysql; will probably be nagios default in future
- stalking: logging all
- often integrated with cacti/rrdtool for trend monitoring?
- monitoring systems (2008)
- nagios: see section
- zenoss (core): zope
- web/daemon/data-tier arch; data in cmdb (cfg), rrd (hist), mysql (events)
- uses snmp (req'd), nagios plugins, zenpacks (cmds, tmpls, graphs)
- all cfg via web ui or api
- tmpls; production states for servers; alert severities; locations
- integrates events and trends; nice/confusing gui
- scalable server
- missing commercial features in zenoss enterprise
- zabbix
- lightweight agent/php/db-tier arch
- simple checks; agents; snmp; external/internal/aggregated checks
- auto discovery (agent based)
- fully integrated alerting, reporting, trending
- create own screens w tmpls; correlation of graphs
- cons: cumbersome config; reporting not great; no escalation
- hyperichq
- java; pg
- focus on app internals (mysql, pg, jboss); great for debugging
- integrated with opennms; similar frameworks; complementary
- auto-detection; best graphing; easy config; nice GUI
- not for typical lamp shop; lots of features in commercial version
- opsview enterprise
- set of nagios exts for scaling, web (catalyst), mysql
- HA/LB monitoring; multi slaves, single master
- nagvis, rrdtool for trending
- opennms: once nagios' only contender; j2ee; focus on network; snmp + nagios
plugins
- groundworks: sucks
- hobbit: sucks
- from <http://www.slideshare.net/tomdc/open-source-monitoring-tools-shootout>
- nice feature matrix
- scipy
- scikits
- openopt: bunch of optimization
- sqlite
- be careful, int != integer:
sqlite> create table a (a integer primary key, b integer);
sqlite> insert into a (b) values (0);
sqlite> select * from a;
1|0
sqlite> create table b (a int primary key, b integer);
sqlite> insert into b (b) values (0);
sqlite> select * from b;
|0
- R
- packages
- misc
- Rserve
- foreach
- parallel
- lubridate
- data
- RPostgreSQL
- RGoogleDocs
- sqldf: SQL on DFs
- plyr: general bulk manipulation utils
- reshape2: pivot table-like cubing/rollups
- data.table: a better data.frame
- xtc: builds on zoo; "unifies" the various time series pkgs
- model selection
- leaps: stepwise/exhaustive, AIC/BIC/R2/...
- lars: LARS, LASSO, fwd stagewise
- OLS only
- faster than glmnet for small/sparse/wide problems:
<http://stats.stackexchange.com/questions/7057/glmnet-or-lars-for-computing-lasso-solutions>
- lasso2: L1 for GLM
- clustering
- mclust: mixture modeling, model-based clustering
- modeling
- randomForest
- survival
- rjags
- mgcv: smoothers for GLMs
- caret: data splitting, preprocessing, model tuning using resampling,
variable importance estimation
- train: parallel training procedure that takes care of a lot of stuff
for you
- knn
- ROCR
- prediction
- performance
- kernlab
- e1071: misc functions from dept stats, TU Wien
- NB, SVM, perms/combs, shortest paths, distances, LCA, lots more
- tm: term document matrices
- irlba: SVD for big data
- arm: bayesglm
- assorted
- MASS: companion to modern applied statistics with s
- rlm: robust regression
- boxcox
- stepAIC: fuller version of `step`
- numeric
- quadprog
- eval
- verification: fast roc.area & other measures
- data
- discretization: mdlp same as orange.EntropyDiscretization
- viz
- ggplot2
- ggExtra: `install.packages("ggExtra", repos="http://R-Forge.R-project.org")`
- rggobi
- directlabels
- lattice (splom), car (spm): scatterplot matrices (see ggplot2)
- maps, geosphere
- popbio
- logi.hist.plot: viz logistic regression results
- hexbin
- HH
- ci.plot: confidence interval plot
<http://princeofslides.blogspot.com/2011/02/sab-r-metrics-basic-applied-regression.html>
- finance
- quantmod
- quantstrat
- BurStFin
- vocab
- modeling facilities
- expressions: `y ~ a + b + c:d + e*f - b` (`c:d` means interaction,
`e*f` means `e + f + e:f`)
- linear models: lm, glm
- clustering: kmeans
- optimization: optim, nlm
- ANOVA: aov, anova
- descriptive tools: table, summary
- gbm: gradient boosting machine
- nonlinear models: loess, scatter.smooth
- model selection: step
- eval: rstandard, rstudent, AIC, BIC
- misc: predict, resid, coef, confint, cooks.distance, diffits,
influence.measures
- stats
- cor, cov, var, mean, median
- summary, table
- system
- install.packages, update.packages
- setwd, read.csv
- options, sessionInfo
- trace, traceback, stop, stopifnot
- png, dev.off, par, dev.new
- manipulation
- transform, head, tail, subset, cbind, rbind, colnames
- abline, lines
- str, dput
- all.equal, identical
- new.env, with, within, search, ls, attach, detach, get, assign, `<<-`,
`<-`
- sort
- na.*
- match, %in% (subset), is.element, setdiff, merge
- paste
- apply, lapply, sapply, tapply
- rbind.fill
- cut
- control
- invisible
- syntax tricks
- use backticks for non-standard names
- to run an R file from cmd line
- `R --vanilla "--args ${@:2}" < $1 #>/dev/null`
- will be replaced with `Rscript`
- GUIs: RStudio beats all the rest
- pitfalls
- don't pass in logicals as `y` to `train`; use `factor(ifelse(x, 'pos', 'neg'))`
- RF can only have categorical variables with <=32 factors
<http://r.789695.n4.nabble.com/randomForest-can-not-handle-categorical-predictors-with-more-than-32-categories-td3036871.html>
- gripes
- syntax: can never tell when something will be evaluated, when i'm writing
symbols or formulae or expressions, etc.
- inconsistency: packages use model formulae, vectors/matrices, and/or
data.frames
- hard to find packages, determine best packages, etc.
- memory management/fragmentation
- caret
- can't dynamically choose params (eg glmnet lambda, randomForest mtry)
- smart cards
- form factors: credit card, mini-SIM (most common), micro-SIM
- communication w PCs
- PC/SC (personal computer/smart card): serial protocol; in Win; PC/SC Lite
for Linux
- CCID, ICCD: USB protocols
- ubuntu
- `-dbgsym` packages have debug symbols for all binaries
- user adm is for viewing system logs; user admin is for sudo
- debian
- apt
- /var/run/reboot-required: presence means need to reboot
- LDAP
- hierarchical directory of entries of KV attr pairs (defined by schema)
- schemas define _object classes_; published at a base DN in
subschemaSubentry _operational attr_ (system attr)
- obj ID is _distinguished name (DN)_
- topmost levels usu. reflect DNS (eg `dc=example,dc=org`)
- directory system agent (DSA): an LDAP server; default port 389
- can refer to other servers
- request types
- start TLS
- bind: authenticate w user DN and passwd, specify version (normally LDAPv3)
- search: base obj, scope, filter, deref aliases?, attrs (projection),
size/time limits, types only?
- compare: check DN has attr pair
- add/delete/modify: takes DN and list of attrs
- extended ops: eg cancel, passwd modify
- abandon: abort an operation named by a msg ID; unfortunately no response,
so cancel extension was introduced
- unbind: abandon ops, close conn
- active directory
- LDAP, modified KRB, DNS, SSO, app roaming storage, admin/policy/deploy
SW/etc
- AD includes 1+ domains, each w 1+ _domain controller (DC)_ (can scale)
- AD app mode (ADAM): lightweight impl on 2k3/XP; same but no need for domain
- primary DC, backup DC, or domain member: older NT model
- now 1+ equal peer DCs; multi-master repl (configurable sync/async/etc)
- introduced in win2k
- samba: impls SMB/CIFS and domain membership; v4 can serve as AD DC (future)
- NBT, SMB/CIFS, DCE/RPC, MSRPC, network neighborhood, WINS server called
NBNS, SAM, LSA, SPOOLSS, NTLM, AD logon
- `smbd` and `nmbd`
- pylons/paster
- paster
- `sys.exit(load_entry_point('PasteScript','console_scripts','paster'))`
- main code calls `get_commands` to find all paster plugins (uses
entrypoints?)
- `--plugin` for extensibility
- pylons
- `appconfig` loads config
- `loadapp` loads config and creates app
- `loadserver` loads config and creates server
- all above take URLs like `config:development.ini`
- `paster setup-app development.ini`
- `[app:main]` in .ini points to egg for your app ("myapp")
- myapp's `paste.app_install` entrypoint is `pylons.util.PylonsInstaller`
- `Installer` objects run `setup_app` (or `setup_config`) in `websetup`
- passes paste's `appconfig()` into `setup_app`
- `paster serve development.ini`
- serve command calls `loadapp` and `loadserver`
- `loadapp` uses entrypoint `paste.app_factory`, points to
`myapp.config.middleware`
- `middleware` is passed `global_conf` (`[DEFAULT]` in .ini) and `app_conf`
(`[app:main]` in .ini)
- calls `load_environment`: creates/returns `pylons.config =
pylons.configuration.PylonsConfig`
- creates `PylonsApp` as stack of middleware initialized by `config`
- the config is then registered into `pylons.config` by `PylonsApp` via
`paste.registry.register` (pushed onto a stacked proxy object)
- config has local_conf and global_conf sub-dicts
- <http://pylonsbook.com/en/1.1/pylons-internal-architecture.html>
- <http://wiki.pylonshq.com/display/pylonscookbook/Pylons+Execution+Analysis+0.10>
- python packages/eggs
- distributed/pip: successor to setuptools/easy_install
- pkg_resources: most of the meat is in this module
- entry points: simple global registry for all modules to declare what
interfaces they provide (and for others to find them by these interfaces)
- egg metadata stored in ./package.egg_info/ or
site-packages/package.egg/EGG-INFO
- net booting, old to new: etherboot (95), intel PXE (99, TFTP), gPXE (SAN)
- android
- processes can be shared by multiple apps
- on backgrounding, apps are signaled to save their state
- OS will kill bg procs under pressure
- philosophy: don't distinguish btwn background/stopped apps
- broadcast receivers: given 10 seconds to handle broadcast events like alarms and arrived-at-location
- services: longer-running bg ops
- email agents
- MUA: client for reading/composing messages, eg mutt
- MSA: recv from MUA, send to MTA; most MTAs also MSAs
- uses: correct errors (eg missing date/msg-id/domain), simplify MTA
policies (refuse mail to non-locals, stricter spam setting)
- MTA aka relay: xfer msgs to/from this host; uses SMTP
- exim is officially supported on ubuntu; also sendmail (older, famously
insecure)
- each MTA adds `Received` hdr
- major MTAs (monolithic = less secure)
- sendmail: popular, once-insecure, monolithic
- exim: monolithic, ok security record, sendmail drop-in
- postfix: modular, secure (by sec guy), performant, sendmail drop-in,
easy
- qmail: secure (by djb), updated in 99, modular, weird to use
- <http://shearer.org/MTA_Comparison>
- MTAs have configurable policies/processing; eg postfix can spawn a cmd,
pipe to a persistent cmd, etc
- port 587 used for MUA-MSA; port 25 used for MTA-MTA; most servers just
use 25
- MDA: recv msgs to local user, store in inbox; eg procmail (old; most MTAs
are also MDAs)
- MRA: retrieve via POP, IMAP, etc; eg fetchmail, getmail (python successor)
- outlook, thunderbird, etc: MUAs that also have some MSA, MDA, MRA
- email configuration
- /etc/mailname: the domain of this machine
- eg cron jobs are from/to this domain
- this is what postfix should set as mydestination
- fail2ban: scans logs and bans IPs w too many failed attempts
- update iptables FW or tcp wrapper's hosts.deny
- php
- `php -a`: if built w readline support, then behave as console
- ruby web app infrastructure
- monit: monitors mongrel processes (and unicorn masters)
- mongrel: ruby web server
- unicorn: request queue for server load balancing; based on mongrel
- features 0-downtime app deployment (incremental rollout)
- stormcloud: monitors unicorn workers
- xen
- dom0: hypervisor kernel
- domU: guest kernel
- hadoop
- containers (sources/sinks)
- files
- SequenceFile: provides len-delimiting btwn keys/values and btwn records
- uncompressed, record-compressed, block-compressed
- built-in compression codecs: bz2, gzip, default (?)
- InputFormat: FileInputFormat:
- SequenceFileInputFormat: uses SequenceFileRecordReader
- TextInputFormat: uses LineRecordReader
- KeyValueInputFormat: tab-separated lines
- serialization
- Writable: WritableComparable:
- hadoop records aka Jute
- used in RPC
- primitives: byte, bool, int, long, float, double, ustring, buffer
- buffer (BytesWritable): use this for eg protobufs
- composites: record, vector, map
- Record: abstract; for generated classes
- DDL (link.jr) to `rcc -l C++` or `Java`
module links {
class Link {
ustring URL;
boolean isRelative;
ustring anchorText;
};
}
- encodings: binary, CSV, XML
- avro
- hadoop
- used to require keys/values to be Writables, but now uses Serializers
and class factories
- there's WritableSerialization
- old
- RecordReader: also had DBRecordReader
- <http://code.google.com/p/thrift-protobuf-compare/wiki/Benchmarking>
- <http://archive.cloudera.com/docs/>
- setup: <y_z.scripts.mit.edu/wp/2010/01/20/no-nonsense-standalone-hadoop-and-dumbo-on-ubuntu/>
- scala
- vocab
- `@tailrec`
- `@cps`
- util.control.Exception.catching
- Stream.continually
- `breakable`: somehow faster than simple Exceptions
- things i don't like
- complexity, almost entirely in the type system which has many interacting
concerns
- <http://stackoverflow.com/questions/1715681/scala-2-8-breakout>
- open questions
- <http://stackoverflow.com/questions/5544536/type-inference-on-set-failing>
- <https://groups.google.com/forum/#!topic/scala-user/NwhaQ6U0wy0>
- <http://stackoverflow.com/questions/6448444/missing-parameter-type-in-overloaded-generic-method-taking-a-function-argument>
- <http://stackoverflow.com/questions/7830731/parameter-type-in-structural-refinement-may-not-refer-to-an-abstract-type-defin>
- <http://stackoverflow.com/questions/7829765/scala-view-bounds-that-work-with-subtypes>
- <http://stackoverflow.com/questions/6849994/understanding-the-interaction-in-scala-between-self-types-and-type-bounds>
- <http://stackoverflow.com/questions/6849137/scala-type-parameter-bound-error-in-subclass-but-not-superclass>
- collections library was hard to get right (original sucked), and is
hard to comprehend
- hopefully some clever and motivated minds will make things better in
the future
- sbt
- too hard to get started with (docs are better now)
- still too complex to understand and extend
- Mark Harrah is difficult and uncommunicative
- simple things/extensions are cumbersome:
- <http://stackoverflow.com/questions/7134993/how-do-i-run-an-sbt-main-class-from-the-shell-as-normal-command-line-program>
- <http://comments.gmane.org/gmane.comp.lang.scala.simple-build-tool/2102>
- collections branching off from TraversableOnce into Iterator and Seq
- bunch of methods missing for TraversableOnce
- methods missing unnecessarily between Iterator/Seq, e.g. distinct
- syntax
- need acute awareness of what we're dealing with
- python 3 gets this right: produce iterators by default, since
everything else can be efficient build on top of them
- in general, self-type-returning collections are not a good default
- `for (x <- xs) ...` may or may not produce an intermediate list,
depending on the type of `xs`; better remember to always use
`.iterator`
- see also the need for `breakOut`
- used to be a die-hard user of Streams, but now almost never use them
- for comprehensions and lack of generators:
def splitBy(f: (A,A) => Boolean) = {
var ys = mut.Buffer[A]()
for {
Seq(Some(a),b) <- padIter(xs) sliding 2
if { ys += a; b == None || f(a,b.get) }
} yield {
val oldYs = ys
ys = mut.Buffer[A]()
oldYs
}
vs.:
group = []
for a,b in sliding(xs, 2):
group.append(a)
if a != b
yield group
group = []
- for comprehensions vs python generators/yield
- for comprehension syntax is its own little world
- nits
- `:` at the end of an operator method means `this` is the RHS
- why aren't all functions curried? why does currying need to be specified in advance?
- abuse of operators: `/:` is unnecessary
- why both classes and case classes? either new Foo or Foo? just have one.
- needs to decide whether it wants to integrate with java or not. (don't!)
- `:_*` syntax
- singleton objects are promoted everywhere and are also the main way a lot of
tricks in scala are pulled off.
- why are `for`, `if`, `while`, etc. special in the language? ideally should be
implemented in the language.
- breakOut
- List.map returns List instead of iterator? (cf. Python)
- language is too expansive
- Option *and* null?
- XML literals, yet no string interpolation
- more serious
- too many choices all the time: do I want a function, a view method, etc.
- libraries aren't as mature (complete, battle tested, documented)
- lots of compat breaks
- IDE is a WIP
- syntax just doesn't flow very smoothly
- the infix operator syntax always needs to be wrapped in parens if you
then want to apply a unary postfix method, discouraging its use if
you want to avoid having to jump around to place parens
- the no-() method invocation fails when you want to call an apply on
the returned result but the orig method has an implicit param that
gets in the way
- longer term
- no javascript backend
- no native backend
- no `dynamic` - `Dynamic` will come close but not entirely and it's not
here yet
- html
- link rel canonical: the canonical URL for this page; useful for eg search
engines, sharing, bookmarking, etc
- glibc
- buffering: fully buffered (default), line buffered (tty), unbuffered
- see `setbuf(3)`, `stdbuf(1)`
- pdf
- <http://www.adobe.com/devnet/livecycle/articles/lc_pdf_overview_format.pdf>
- <http://www.mactech.com/articles/mactech/Vol.15/15.09/PDFIntro/>
- disks
- sector: 512 B
- partition: physical region on disk; serves as a storage container
- windows PC: first partition starts at sector 63
- volume: OS abstraction for storage container
- usu just a single partition
- logical volume managers (LVMs) can create virtual volumes; eg concat or
stripe partitions
- lvm
- features: live PV adds/removes; resize partitions; snapshots; span many
disks; RAID0/1
- _physical volumes (PVs)_ are real partitions or whole disks
- _volume groups (VGs)_ of PVs are logical disks
- _logical volumes (LVs)_ are partitions of VGs
- can resize LVs; just first resize FS to be same/smaller
- reiserfs has fast resizing, unlike ext3
- lvm2: read-write snapshots; fragile, don't be too fancy
- default extent ('chunk') size 4MB
- keeps metadata header at start of each PV; each PV has UUID
- complete copy of entire VG layout, incl other PV UUIDs
- LVM implemented in terms of the _device mapper_; simplifies LVM code; all
in user space
- /boot can't be in LVM; / not recommended
- recommend LVM above RAID
- <http://www.ntlug.org/Articles/LVM>
- gdb
- mostly implemented using ptrace
- athena
- `system:htaccess.mit`: the web server group
- random vocab
- ILP32: int, long, ptr are 32 bits
- LP64: long, ptr are 64 bits
- other RDBMSs
- main-memory hsqldb: only `read uncommitted`
- mysql
- innodb stores data in PK order
- explain/optimization
- select types
- simple: no union or subqueries
- primary: outermost select
- subquery
- dependent subquery: correlated subquery
- derived: subquery in `from`
- uncacheable subquery: must re-evaluate for each outer tuple
- join/access types, best to worst
- const: index lookup yielding exactly 1 row; can turn the result of this
index lookup into constant
SELECT * FROM tbl_name WHERE primary_key=1;
SELECT * FROM tbl_name
WHERE primary_key_part1=1 AND primary_key_part2=2;
- eq_ref: ref where the inner table yields exactly 1 row
SELECT * FROM ref_table,other_table
WHERE ref_table.key_column=other_table.column;
SELECT * FROM ref_table,other_table
WHERE ref_table.key_column_part1=other_table.column
AND ref_table.key_column_part2=1;
- ref: index join
SELECT * FROM ref_table WHERE key_column=expr;
SELECT * FROM ref_table,other_table
WHERE ref_table.key_column=other_table.column;
SELECT * FROM ref_table,other_table
WHERE ref_table.key_column_part1=other_table.column
AND ref_table.key_column_part2=1;
- index_merge: use multiple indexes
- unique_subquery: non-dependent (non-correlated) subquery using unique
index; turn index lookups into constants
- index_subquery: non-correlated subquery using non-unique index; turn
index lookups into constants
- range: use index to select a range;
- index: index scan (per outer row); only if predicate and selected
columns are covered by index
- all: full table scan (per outer row)
- columns
- key: the key used by this join/access type
- ref: which outer columns are compared against `key`
- innodb
- perf
- trx_commit, binlog, sync_bin
<http://www.mysqlperformanceblog.com/2010/02/28/maximal-write-througput-in-mysql/>
- implements MVCC; has true serializability bc of shared index locks,
unlike PG (and future locks or gap locks can be taken on these)
- each row has 3 new fields
- 7-byte ID of last updating (or deleting) txn
- 7-byte ptr to undo log record (undo log is called the _rollback
segment_)
- 6-byte monotonically increasing row ID
- <http://dev.mysql.com/doc/refman/5.0/en/innodb-multi-versioning.html>
- storage format
- if table has int primary key, use that as tuple id; otherwise add tuple
id for each tuple
- var-len cols
- compact/redundant ("antelope"): up to 768B in record (for prefix
indexes), rest in overflow page
- at least 2 rows + some metadata must fit in each 16KB page, so
limit for whole row is ~8KB
- dynamic/compressed ("barracuda"): 20B ptr whenever not fit; prefix
index can be built separately
- each value has exclusive overflow pages; no sharing
- monitoring/mgmt
- `show engine innodb status`
- `show processlist`
- `kill 3` kills process 3 in processlist
- clustering
- built-in async stmt-/row-based replication
- ship stmts/rows/mixed via binary log
- log after update completion but before lock release/commit to ensure
log is in execution order
- mixed: rows automatically used when non-deterministic queries executed
- single thread on slave replays log updates
- only innodb synchronizes w the binary log
- row-based supported by innodb in some cases
- auto-inc fields lock table for non-simple inserts
- setup/resync of slave is a bit complicated
- manual monitoring/failover with `stop slave; reset master` on
promotee and `change master to` on new slaves
- first, wait for `stop slave io_thread; show processlist` to say "Has
read all relay log"
- more details: <http://dev.mysql.com/doc/refman/5.5/en/replication-solutions-switch.html>
- google's 5.5 patches add semisync replication
- master confirms receipt & logging of update by at least one slave
- only adds wait after commit completes; since commit actually happened,
master & slave are actually out of sync (eg if crash during wait,
committed txn may not have made it to slave)
- mysql cluster: ndb with mysql
- ndb is also some sort of standalone clustering solution
- ndb supports scalable sync replication & partitioning using 2PC
- in-mem; async disk log of redo records and chkpts (2s log write period)
- much slower than standard replication
- can replicate btwn ndb clusters or btwn ndb & other engines via
standard replication
- tungsten replicator
- simple setup, stmt-based, sync, global txids, cross-dbms
- supports mysql, oracle, jdbc; future: postgresql, others
- <http://www.wikivs.com/wiki/MySQL_vs_PostgreSQL#Replication_and_High_Availability>
- examples of why mysql blows vs postgresql
- ref: <http://now.eloqua.com/e/er.aspx?s=400&lid=235&elq=008b3ac42ba641fb851200ffd6633d90>
- <http://news.ycombinator.com/item?id=2176062>
> "InnoDB *is* still broken...Just last week we had to drop/re-create an
> InnoDB-table in one project because it would not allow to add an index
> anymore, no matter what we tried...`Mysql::Error: Incorrect key file
> for table 'foo'; try to repair it: CREATE INDEX [...]` "
- pg now has async replication; that was mysql's only strong pt against pg
- neither has synchronous replication, but pg's is coming
- semisync still unsafe; master waits *after* actually committing
locally, instead of blocking the commit (as it should)
- recovery is complex
- no group commit <http://www.mysqlperformanceblog.com/2011/07/13/testing-the-group-commit-fix/>
- crappy concurrency, >3 sucks vs pg <http://spyced.blogspot.com/2006/12/benchmark-postgresql-beats-stuffing.html>
- multiple storage engines has always restricted progress:
<http://www.mysqlperformanceblog.com/2010/05/08/the-doom-of-multiple-storage-engines/>
- <http://stackoverflow.com/questions/324935/mysql-with-clause>
- <http://bugs.mysql.com/bug.php?id=10327>
- postgresql also supported multiple storage engines in 80s; concentrated
on one
- mysql's engine arch is poor; bad to opt/plan/tune
- only recently
- per-stmt triggers
- procedural lang support
- only own internal auth system; pg supports a wide array
- no referential integrity
- no constraints (`check`)
- no sort merge join (certainly no map-reduce contender), let alone
hash-join
- postgresql has more supple `alter table` impl
- mysql doesn't support ASC/DESC clauses for indexes
<http://explainextended.com/2010/11/02/mixed-ascdesc-sorting-in-mysql/>
- <http://www.dbms2.com/2008/07/10/how-is-mysqls-join-performance-these-days/>
- <http://www.mysqlperformanceblog.com/2006/06/09/why-mysql-could-be-slow-with-large-tables/>
- optimizer only recently started working properly with certain subqueries
- auto_increment jumps up to next power of 2 but inconsistently across
versions (platforms?); this is undocumented
- crappy errors: "Incorrect key file for table 'stock'; try to repair it"
on "alter table stock add constraint pk_stock primary key (s_w_id,
s_i_id);" where `stock` is in InnoDB (which has no "repair table") means
I have no `/tmp` space (no Google answers)
- crappy EXPLAIN output
- innodb auto-extends ibdata1 file; only way to trim (garbage collect) is
dumping and loading
- scoping is broken
mysql> create table t(a int, b int);
Query OK, 0 rows affected (3.30 sec)
mysql> select a, (select count(*) from (select b from t where a = u.a group by b) v) from t u;
ERROR 1054 (42S22): Unknown column 'u.a' in 'where clause'
- optimizer sucks
mysql> SELECT distinct t1.tableid FROM transactionlog t1 JOIN transactionlog t2 ON t1.tableid=t2.tableid AND t1.tupleid=t2.tupleid AND t1.partition!=t2.partition LIMIT 100;
+-----------+
| tableid |
+-----------+
| warehouse |
| district |
+-----------+
2 rows in set (48 min 15.52 sec)
mysql> SELECT distinct tableid FROM transactionlog GROUP BY tableid,tupleid HAVING count(distinct partition)>1;
+-----------+
| tableid |
+-----------+
| district |
| warehouse |
+-----------+
2 rows in set (36.33 sec)
- tips
- `set global general_log = 'ON'` or `'OFF'`
- `set session sql_log_off = 1`
- `load data infile` 20x faster `insert`; else, use txns, `insert delayed`,
multi-valued `insert`
- `--safe-updates`/`--i-am-a-dummy`: guard against mass deletes/updates
that are missing `where` clauses
- grant tables: user, db, host, tables_priv, columns_priv
- created by mysql_upgrade and mysql_install_db
- "engine" is newer term for "type"
- `create table t (i int) engine=innodb`
- `--default-storage-engine=innodb`
- built-ins
- myisam: non-txn'l; default; fast
- memory: non-txn'l; non-persistent
- merge: non-txn'l; read-only
- merges identical myisam's into 1 logical table
- insert into physical table; query from merged
- innodb: txn'l; foreign key integrity
- bdb: txn'l; removed in 5.1
- ndbcluster: for mysql cluster; partitioned tables
- archive: large data, no indexes, small footprint (compressed)
- example: stub dev reference
- csv: dev example
- blackhole: takes writes but reads are empty
- DB tools
- mysql2pgsql: pretty good
- EnterpriseDB Migration Wizard: migrates mysql to postgresql
- failed to handle basic `enum` type
Source database connectivity info...
conn =jdbc:mysql://wrench.csail.mit.edu:3307/tpcc
user =remotecarlo
password=******
Target database connectivity info...
conn =jdbc:edb://localhost:5432/tpcc
user =yang
password=******
Importing mysql schema...
Dropping Schema: tpcc
Creating Schema...tpcc
Migrating Tables for Schema tpcc: 'metarelcloud_transactionlog','metarelcloud_graph','metarelcloud_graphsupport'
The data type enum is not handled in Column querytype of Table metarelcloud_transactionlog
Creating Tables...
Creating Table: tpcc.metarelcloud_transactionlog
Error Creating Table metarelcloud_transactionlog:ERROR: type "enum" does not exist
Creating Table: tpcc.metarelcloud_graph
Creating Table: tpcc.metarelcloud_graphsupport
Created 2 tables.
Loading Table Data in 8 MB batches...
Loading Table: metarelcloud_graph ...
Table Data Load Summary: Total Time(s): 0.014 Total Rows: 0
Loading Table: metarelcloud_graphsupport ...
Table Data Load Summary: Total Time(s): 0.008 Total Rows: 0
Data Load Summary: Total Time (sec): 0.022 Total Rows: 0 Total Size(MB): 0.0
Creating Constraint: PRIMARY
Error Creating Constraint PRIMARY
Creating Constraint: PRIMARY
Creating Constraint: PRIMARY
Creating Index: transactionid
Error Creating Index transactionid: ERROR: relation "metarelcloud_transactionlog" does not exist
Creating Index: tableid
Error Creating Index tableid: ERROR: relation "metarelcloud_transactionlog" does not exist
Creating Index: nodeid
Error Creating Index nodeid: ERROR: relation "metarelcloud_transactionlog" does not exist
One or more schema objects could not be imported during the migration process. Please review the migration output for more details.
- postgresql
- PG8.3 added spread chkpts; fsync at end
- limitations/annoyances
- count(*) always scans all rows
- indexed fields have ~2.5K size limit
- clustering
- built-in async streaming record-based replication in 9.0; sync in 9.1
- warm standby: can't be queried; hot standby: can be queried (limited)
- manual failover
- manual recovery
- <http://www.postgresql.org/docs/current/static/continuous-archiving.html>
- <http://developer.postgresql.org/pgdocs/postgres/functions-admin.html>
- <http://archives.postgresql.org/pgsql-general/2010-08/msg00264.php>
- <http://momjian.us/main/writings/pgsql/hot_streaming_rep.pdf>
- pgpool-II
- has had async WAL-shipping replication, but this ships WAL segments
instead of records
- only ships archived WAL segments; controlled with eg `archive_timeout`
- lag is on order of minutes; uses commands like `scp` or `cp`
- to check for stale reads, can compare `pg_current_xlog_location()` and
`pg_last_xlog_replay_location()`
- slony: built by jan weick, postgresql core team member; main option
- much slower/uses more resources than mysql built-in replication
- $O(n^2)$ communication costs
- sql-/trigger-based replication
- auto failover
- 2PC: `prepare transaction`, `commit prepared`, `rollback prepared`
- pgpool-II
- synchronous replication, auto failover, connection pooling/load
balancing, online recovery
- sync replication: 30% write overhead, avoid non-pure queries
- 9.0 async replication: provides auto failover, conn pooling, load bal,
etc.
- pgbouncer: simple, lightweight connection and transaction pooling
- skytools
- pgq: efficient transactional queue
- londiste: async replication via pgq; cross-version support
- plproxy: partitioning/rpc
- WAL
- segments are separate 16MB files
- checkpoints taken by default every 3 WAL segments or 5m
- has async api
- monitoring/introspection
- `select * from pg_stat_activity`: see current queries/activity
- `select * from pg_locks where not granted`: see locks
- rules: query rewriting, eg let you insert/update a view and redirect to
underlying table(s)
- triggers: executed in-transaction
- notify/listen: messaging to clients
- sent only after commit