- Knowledge set
- SNAKE principle
- scenario: case / interface
- necessary: constraints/ hypothesis
- application: service/ algorithm
- shortURL insert
- longURL find
- kilobit: data size
- evolve: how to improve
- jargon
- daily active user
- operation per person
- Query per second
- AngularJS
- 0. 在body层声明ng-app
"<body ng-app="tinyurlApp">"
- 1. 通过ng-view,实现SPA,不同内容显示在主框架内
"<div class="container">
<div class="ng-view"></div>
</div>
"
- 好处1: 外框架代码共用
- 好处2: 不同view,分别对待
- 2. 通过在JS文件中声明ng-Route模块,决定 URL - View - Controller 的对应关系
"一个 View 对应 一个 Controller,代码如下:
app.config(function($routeProvider){
$routeProvider
.when("/",{
templateUrl: "../public/views/home.html",
controller: "homeController"
})
.when("/urls/:shortUrl",{
templateUrl: "../public/views/url.html",
controller: "urlController"
})
;
});"
- URL路径可以使用变量统一获得,变量名字前边加 “ :”
".when("/urls/:shortUrl",{
templateUrl: "../public/views/url.html",
controller: "urlController"
})"
- Node.js
- V8 engine Chrome based on C. All statement be converted into C , so can handle things other that just used for browser
- 1. Event driven:
- 不是按照程序语句顺序执行,而是按照event触发顺序执行
- event listeners + event emitter,不同于函数定义、函数调用。这里将所有event handler 存在一个地方,统一由event listener听取event,进行event与handler的分配.
- 2. 异步 IO
- Node.js比普通的php要好,为啥?
- 因为普通的如果有请求文件或者网络时,php会停止接收新请求,直到上一个io完成。
- 然而nodejs就不会将后续请求block,而是将其请求IO 传给并行IO处理器,由其统一处理IO
- CPU密集型,IO密集型
- PHP如果只有CPU请求,那会更好。它会不断的开辟新线程进行请求处理,。但如果某个线程中,有IO 请求的话,他就会将该线程的IO 处理完,再继续处理后续请求。但是会有一种情况:如果某时间IO 请求数量很多的话,所有的线程都在处理IO ,以至于新的请求没有办法被接受,这就block了。
- Nodejs,统一用一个进程单独接受请求。然后将含有IO请求的子任务统一交给多线程的IO处理模块。虽然看似多线程,但内部是单线程执行cpu request的。它对于CPU密集型的请求很差劲,因为他只用一个线程处理cpu,腾出了多个线程处理IO
- RESTful API
- insert
- POST /urls
- 发送数据
"data:{
longURL
}"
- 返回数据
"方法一:
response:{
id:
longURL:
}
方法二:
response:{
id:
longURL:
shortURL:
}"
- Lookup
- 方法一: GET /urls/{id}
- 方法二: GET/urls? shortURL = xxx
- best way:
- Semantic Versioning
- breaking change + new feature + bug fix
"incompatible API changes + add backwards-compatible functionality + make backwards-compatible bug fix"
- ^ ~ >=
- ^ 自动升级到 new feature
- ~ 自动升级到 new bug-fix
- Bootstrap ( responsive web design )
- multi-device
- desktop / laptop / tablet/ smart phone
- 基于LESS
- 重用CSS代码,可以定义变量
- -webkit-box-shadow: safari引擎
- Gird System
- 先分行(无限)
- 后划分列(12 列)
- col-xs-12
- col-sm-6
- col-md-8
- table
- table-striped
- table-hover
- form
- form-inline
- button
- <a>
- <button>
- <input type='button'>
- <input type='submit'>
- color: button-alert
- componet: glyphicons
- Font awesome( another website)
- dropdown
- JavaScript
- pills
- tabs
- modals(弹窗)
- accordin
- crasoul
- Popular frontend framework
"todoMVC - 学习前段"
- react
- 仅仅针对前段View
- component-based architecture
- 某个变量元素增加后,避免重新生成新html文件。而是仅仅增加必要的dom component
- angular2
- component-based architecturce
- ember
- convention not configuration
- 提供了很多的模块,架构笨重
- backbone.js
- 极简,不笨重
- 难学
- Meteor
- 前段+后端
- 前端显示、后端数据库直接绑定
- Restful API不适用
- react结合较好
- MongoDB
- document database
"JSON type
{
"id" : 111,
"name" : "zhang wen "
}"
- 内存中直接按照String存储
- collections of documents
- how to save data?
- with padding
- but lead to fragment
- how to reduce fragment?
- 单人使用硬盘:padding size is (next power of 2 - current size)
"current is 40K, padding is 24k = 64"
- 多人使用硬盘:
- initial size: 2
- next allocated size: 4/ 8/ 16/ 32 .. 2G
- BSON
- 加入了数据块的字符长度,在检查到Id不是desired时候,很方便跳过该长度的数据块,到下一个
- how to read faster with fragments:
- add pointer to next data set
- how to find index faster?
- Create Index
- B-tree index, ( similar to BST )
- how to accelerate read & write?
- utilize memory as cach
- add log in case of crash
"因为写log硬盘与写data硬盘速度不一样,写data硬盘需要查找, log直接在后边写"
- Redis
- In-memory database
- Single-thread, Event-loop
- 大学食堂例子
- multi-thread
"10个厨师轮流使用锅"
- single-thread
- 如何防止多个线程同时读写同一个资源引起错误:加锁
- but,加锁是限制QPS的主要原因,因为CPU以及网速都没问题
- 不用多线程!只要单线程顺序执行
- Cache
"App ---- Cache ---- Database"
- 1. 一切都是cache;一切都可以实现cache
- replacement(替换策略)
- LRU: least recent use
"double linkedlist"
- Pre-load
"启动之前,先将数据库中数据load到cache中"
- Message Broker
"publish & subscribe"
- Nginx
- server types - big picture
- client tier(angularJS)
- web tier (nginx server)
- business logic tier(node.js)
- database (mongoDB)
- Reverse Proxy
- (正向)代理forward proxy:
- 在client与internet之间,替client访问internet的计算机
- 路由器就是一个forward proxy,外网看不到路由器之后的计算机,但是计算机可以看到外边
- 翻墙服务器
- (反向)代理:
- 在server与internet之间,替server接受internet访问的计算机
- 1. Hide Original server
"安全"
- 2. application firewall
- 3. SSL termination
"加密不需要server进行,反向代理可以代为计算以减轻server压力"
- 4. Load balance
- 1. round robin
"optionally weighted"
- 2. least connected
"optionally weighted"
- 3. IP hash
"1~200: server a; 201~400: server b"
- 4. Generic Hash
"based on devices/ client locations"
- 5. Least time (nginx plux)
"optionally weighted"
- 5. Cache
- GET都可以cache
- 即便是动态列表也进行cache,因为实际生活的改变不会比nginx的处理快
"1s <---> 100000 request/s"
- POST不能cache
- 6. Compression(gzip)
"减小网络传输量;用reverse proxy/ browser 的CPU为代价"
- 7. 不宕机进行配置改变/升级
- new出新worker/master 实例。
- old配置:继续处理完接收到的请求
- new配置:接受新的请求
- Docker
"build, ship and run any app, anywhere
http:// dockerhub"
- why need it?
- 1. 向服务器部署代码非常困难
- why?
- 复杂的技术栈
- why again?
- 不同的技术栈之间版本依赖与冲突
- 2. 以前的practice
"App
Tomcat6
Apache
Linux
Machin"
- 为了给客户看,需要将第一个版本程序打包作为镜像安装在客户的machine上。
- 但是后来技术栈可能会有频繁的改动。这样还要将原来的版本删除,再用改动后的技术栈重新安装,麻烦
- Container vs. Virtual Machine
- VM: 占内存
- C:不用再虚拟OS,直接共享原OS,而且共享Bin lib
- Details:
- Union File System
- 多层file system union起来,由上层决定app运行要用哪些文件。
- 只有上层read-write;下层只read。 保证了上层config.json可以rewrite配置,覆盖之前的配置,从而使用不同的文件,使用不同的port
- Docker Component 架构
- docker daemon(后台程序)
- docker client
- docker image
- docker registries
- docker container
- 配置图(docker compose)
" reverse proxy
|
load balancer
|
web app + web app + web app
|
database"
- Work flow:
- 0. consist of 2 parts:
- 本地Docker engine
"用于运行本地镜像,从而运行程序"
- 远程Docker Image
"从远程pull优秀的镜像到本地,供app使用"
- 1. Similar to Operating System/ Virtual Machine
- 拥有自身的运行逻辑
- 拥有自身的file system
- 2. Programmer write/config the Dockfile to declare which platform/ server/ backend framework/ frontend to use
- 3. build up docker container using those declared images.
- first, the docker is running on our machines, and connect itself with some port of our real machine. Thus, our app is keep running in Docker as a software on our real machine
- second, since it is actually a "virtual machine", it has it own port number, we have to specify the mapping between our real machine and the virtual machine
- third, it means we have a virtual machine running our app using specified platform(Linux) / server(nginx) / backend framework(NodeJS) /frontend framework*AngularJS)
- 4. outside visitors browse our web app via IP+Port. Docker
- first, the request is forwarded to docker port, since
- first, it is running as an software on our real machine
- second, the port number of our real machine channels the request to Docker to its mapping port
- 5. In Docker, top layer is running our app, by residing/ depending on its lower layers.
-
- Cassandra
"distributed system"
- 点与点之间如何沟通?
- Gossip Protocol
"一传十,十传百"
- NoSQL
- 关系型数据库
- 持久性:
"日志rollback+备份backup"
- ACID
"atomicity + consistency + isolation + durability"
- 逻辑视图 != 物理视图
- 购物transcation所设计的相关的变量,并没有在物理存储中没有清楚的关联
- 聚合型数据库
- 分类
- key-value
- document database
- column family store
- graph database
- 虽没有显示schema,仍有隐式的schema
- 只按照一种方式聚合(每次都按照customer进行存储)
- how to choose?
- 查询是否只按照一种方式进行
- 拓展( Extension )
- Cache
- LRU, Pre-load
- User Management System
- Real-time Analysis
- websocket
- URL Expiration
- Work flow
- 1. Users come to the app webpage
- Browser send GET request with original URL to server
- Server sends index.html to clients
- Browser get the index.html with "<script src ='angularJS library '> <script src = 'app.js'> <script src = '/public/js/controllers/homeController.js'> in the head tag.
- Browser again sends request for angularJS library/ js controller files to different servers
- when browser encounter 'ng-view', it automatically turn to 'ng-router', ask for its view and controller; Again, sends corresponding GET request to fetch sub-view files
- while browser parse the controller js file, it might do AJAX request to fetch information from server.
- 2. Users use the shortening service
- input longURL in the input box, and click 'submit'
- 'submit' fires the onclick funcion to do 2 things:
- 1. send the longURL to specified URL, ie. POST to localhost/api/v1/urls
"Notice: This url is specifically used for rendering short/long URL"
- 2. jump to new subpage: url.html, using $location.path( "/urls/" + shortURL )
"Notice: this step is executed after app server responds the POST request in the first step with the calculated shortURL "
- app server first receive the POST request with longURL, do 2 things:
- 1. call the urlService to calculate the shortURL, store long/short URL pair in MongoDB
- 2. send back the shortURL to the app webpage.
"NOTICE: the shortURL should consist different prefix compared with the current app URL. Because it is created for outside users. So the shortURL should not prefix with /api/v1/.., but with just '/:shortUrl'"
- after the app webpage getting shortURL, it renders new subview using url.html; Meanwhile, it needs to know the its original longURL for display.
- Again, in the url.html subview, it sends GET request to the app server to get the longURL
"Notice: This step could be simplified by just passing the longURL from the first subview to the second subview creating a service or using the $rootScope"
- Now, in the new subview, the longURL and shortURL pair can be displayed
- 3. Visitors browse the shortURL from outside of the app webpage with 'localhost/:shortUrl'
- the shortURL is just clicked from various browsers/ machines/ from different IP/ countries, at different time
- when shortURL is clicked, it sends GET request 'localhost/:shortUrl' to the app server.
- the app sever call the urlService to get the longURL from the shortURL, and redirect users to the original long URL
- 4. Information statistics
- we want to record the total number of clicks / their browsers/ machine brands/ their IP/ countries/ timestamp. And display them in the app webpage.
- So first we need add more elements in the url.html subview in the frontend; And update those information in the backend.
- 1. updating the info each time the shortURL is clicked. And then store them in the MongoDB
- 2. when app webpage is clicked, make AJAX call to the app server, let it get the info from MongoDB, and then send back to frontend
- since all these clicks directly send GET request to 'localhost/:shortURL', which is handled by redirect.js. Thus, we here need a statistic service to update these info.
- here we need the one of Express modules: Express-useragent to extract the information from the request
- In the statistic service, we use the express-useragent to get the info, and then store them in the MongoDB database with a new schema.
- 5. app server's functions:
- 1st function: handle app webpage request
- 2nd function: handle outside's request
- difference: clients shortURL request vs. app shortURL request
- clients 点击 shortURL时候,nodeJS会根据express的route进行该shortURL的redirect
"【注意】这些shortURL并不是从本网页点击的,而是从外部电脑访问该shortURL的"
- app frontPage请求shortURL时候,其实是给用户一个查看数据的方式,只有他输入原始URL,才能得到相应的shortURL,才能查到相应的访问信息
"【注意】该shortURL是从本网页访问的"
- route has to cover all the cases, including browsers send GET request to get its static resources. Thus, the route has to declare the "/public " path and its corresponding Router.